Niko Beerenwinkel

dblp:58/2558 · DBLP profile ↗
← Back
83ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-0573-6119ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 78 · 8 first-author · 23 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Quantifying uncertainty of predictions from cancer progression models
abstract
MOTIVATION: Cancer progresses through the accumulation of genomic events. Cancer progression models such as Mutual Hazard Networks (MHNs) describe this dynamic, enabling prediction of temporal event positions and patient-specific risks of acquiring mutations. However, current MHN analyses rely on single most likely models and do not quantify the uncertainty inherent to parameter estimation. Assessing forecast stability is essential before using them to anticipate treatment-relevant mutations, adapt targeted therapies, or prioritize monitoring of patients at elevated progression risk. RESULTS: We address a key prerequisite for the responsible clinical use of cancer progression models by making MHN-derived predictions uncertainty-aware. We present a Bayesian framework for MHN that uses Markov Chain Monte Carlo to sample from the posterior distributions of model parameters and derived predictions. For practical use we implemented the Random-Walk Metropolis, Metropolis-Adjusted Langevin Algorithm (MALA), and simplified manifold MALA samplers as part of the existing mhn Python package. Only MALA and smMALA were successful in sampling from MHN posteriors, with MALA performing best. While most MHN parameters and predictions showed low posterior variance, a small subset displayed greater variability across the posterior distribution. This differentiation cannot be obtained from a single most likely model, emphasizing the need for uncertainty quantification, especially in clinical contexts. As an illustrative example, posterior sampling identified a subgroup of STK11$-$, KRAS$+$ lung adenocarcinoma patients with a high predicted short-term risk-with low variance across posterior samples-to develop an STK11 mutation. This subgroup exhibited poorer survival under immunotherapy, resembling patterns observed in STK11+ patients. AVAILABILITY AND IMPLEMENTATION: Our implementation is part of version 1.2.0 of the mhn package (https://github.com/spang-lab/LearnMHN). All analyses including the code to produce all figures in this article can be found under https://github.com/huy29433/MCMC-sampling-for-MHN (https://doi.org/10.5281/zenodo.21160219).
Y. Linda Hu, Simon Pfahler, Andreas Lösch, Stefan Vocht, Stefan Hansch, Kevin Rupp, Niko Beerenwinkel, Tilo Wettig, Rudolf Schill, Rainer Spang
Bioinform.7
2026 Tracking SARS-CoV-2 genomic variants in wastewater sequencing data with LolliPop
abstract
During the COVID-19 pandemic, wastewater-based epidemiology has progressively taken a central role as a pathogen surveillance tool. Tracking viral loads and variant outbreaks in sewage offers advantages over clinical surveillance methods by providing estimates not biased by testing practices and enabling early detection. However, wastewater-based epidemiology poses new computational research questions that need to be solved in order for this approach to be implemented broadly and successfully. Here, we address the variant deconvolution problem, where we aim to estimate the relative abundances of genomic variants from next-generation sequencing data of a mixed wastewater sample. We introduce LolliPop, a computational method to solve the variant deconvolution problem. LolliPop is tailored to wastewater time series sequencing data and applies temporal regularization in the form of a fused ridge penalty. We show that this regularization is equivalent to kernel smoothing and that it makes abundance estimates robust to very high levels of missing data, which is common for wastewater sequencing. We use the bootstrap to produce confidence intervals, and develop analytical standard errors that can produce similar confidence intervals at a fraction of the computational cost. We demonstrate the application of our method to data from the Swiss wastewater surveillance efforts as well as on simulated data.
David Dreifuss, Ivan Topolsky, Pelin Icer Baykal, Niko Beerenwinkel
PLoS Comput. Biol.4
2025 ELLIPSIS: robust quantification of splicing in scRNA-seq
abstract
MOTIVATION: Alternative splicing is a tightly regulated biological process, that due to its cell type specific behavior, calls for analysis at the single cell level. However, quantifying differential splicing in scRNA-seq is challenging due to low and uneven coverage. Hereto, we developed ELLIPSIS, a tool for robust quantification of splicing in scRNA-seq that leverages locally observed read coverage with conservation of flow and intra-cell type similarity properties. Additionally, it is also able to quantify splicing in novel splicing events, which is extremely important in cancer cells where lots of novel splicing events occur. RESULTS: Application of ELLIPSIS to simulated data proves that our method is able to robustly estimate Percent Spliced In values in simulated data, and allows to reliably detect differential splicing between cell types. Using ELLIPSIS on glioblastoma scRNA-seq data, we identified genes that are differentially spliced between cancer cells in the tumor core and infiltrating cancer cells found in peripheral tissue. These genes showed to play a role in a.o. cell migration and motility, cell projection organization, and neuron projection guidance. AVAILABILITY AND IMPLEMENTATION: ELLIPSIS quantification tool: https://github.com/MarchalLab/ELLIPSIS.git.
Marie Van Hecke, Niko Beerenwinkel, Thibault Lootens, Jan Fostier, Robrecht Raedt, Kathleen Marchal
Bioinform.2
2025 Single-cell copy number calling and event history reconstruction
abstract
MOTIVATION: Copy number alterations are driving forces of tumour development and the emergence of intra-tumour heterogeneity. A comprehensive picture of these genomic aberrations is therefore essential for the development of personalised and precise cancer diagnostics and therapies. Single-cell sequencing offers the highest resolution for copy number profiling down to the level of individual cells. Recent high-throughput protocols allow for the processing of hundreds of cells through shallow whole-genome DNA sequencing. The resulting low read-depth data poses substantial statistical and computational challenges to the identification of copy number alterations. RESULTS: We developed SCICoNE, a statistical model and MCMC algorithm tailored to single-cell copy number profiling from shallow whole-genome DNA sequencing data. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells. We show the accuracy of this approach in evaluations on simulated data and demonstrate its practicability in applications to two breast cancer samples from different sequencing protocols. AVAILABILITY AND IMPLEMENTATION: SCICoNE is available at https://github.com/cbg-ethz/SCICoNE.
Jack Kuipers, Mustafa Anil Tuncel, Pedro F. Ferreira, Katharina Jahn 0001, Niko Beerenwinkel
Bioinform.5
2025 Bayesian inference of fitness landscapes via tree-structured branching processes
abstract
MOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree.
Xiang Ge Luo, Jack Kuipers, Kevin Rupp, Koichi Takahashi, Niko Beerenwinkel
Bioinform.5
2024 Overcoming Observation Bias for Cancer Progression Modeling
Rudolf Schill, Maren Klever, Andreas Lösch, Y. Linda Hu, Stefan Vocht, Kevin Rupp, Lars Grasedyck, Rainer Spang, Niko Beerenwinkel
RECOMB9
2024 Oncotree2vec - a method for embedding and clustering of tumor mutation trees
abstract
MOTIVATION: Understanding the genomic heterogeneity of tumors is an important task in computational oncology, especially in the context of finding personalized treatments based on the genetic profile of each patient's tumor. Tumor clustering that takes into account the temporal order of genetic events, as represented by tumor mutation trees, is a powerful approach for grouping together patients with genetically and evolutionarily similar tumors and can provide insights into discovering tumor subtypes, for more accurate clinical diagnosis and prognosis. RESULTS: Here, we propose oncotree2vec, a method for clustering tumor mutation trees by learning vector representations of mutation trees that capture the different relationships between subclones in an unsupervised manner. Learning low-dimensional tree embeddings facilitates the visualization of relations between trees in large cohorts and can be used for downstream analyses, such as deep learning approaches for single-cell multi-omics data integration. We assessed the performance and the usefulness of our method in three simulation studies and on two real datasets: a cohort of 43 trees from six cancer types with different branching patterns corresponding to different modes of spatial tumor evolution and a cohort of 123 AML mutation trees. AVAILABILITY AND IMPLEMENTATION: https://github.com/cbg-ethz/oncotree2vec.
Monica-Andreea Baciu-Dragan, Niko Beerenwinkel
Bioinform.2
2024 Modeling metastatic progression from cross-sectional cancer genomics data
abstract
MOTIVATION: Metastasis formation is a hallmark of cancer lethality. Yet, metastases are generally unobservable during their early stages of dissemination and spread to distant organs. Genomic datasets of matched primary tumors and metastases may offer insights into the underpinnings and the dynamics of metastasis formation. RESULTS: We present metMHN, a cancer progression model designed to deduce the joint progression of primary tumors and metastases using cross-sectional cancer genomics data. The model elucidates the statistical dependencies among genomic events, the formation of metastasis, and the clinical emergence of both primary tumors and their metastatic counterparts. metMHN enables the chronological reconstruction of mutational sequences and facilitates estimation of the timing of metastatic seeding. In a study of nearly 5000 lung adenocarcinomas, metMHN pinpointed TP53 and EGFR as mediators of metastasis formation. Furthermore, the study revealed that post-seeding adaptation is predominantly influenced by frequent copy number alterations. AVAILABILITY AND IMPLEMENTATION: All datasets and code are available on GitHub at https://github.com/cbg-ethz/metMHN.
Kevin Rupp, Andreas Lösch, Y. Linda Hu, Chenxi Nie, Rudolf Schill, Maren Klever, Simon Pfahler, Lars Grasedyck, Tilo Wettig, Niko Beerenwinkel, Rainer Spang
Bioinform.10
2023 Beyond Normal: On the Evaluation of Mutual Information Estimators
abstract
Mutual information is a general statistical dependency measure which has found applications in representation learning, causality, domain generalization and computational biology. However, mutual information estimators are typically evaluated on simple families of probability distributions, namely multivariate normal distribution and selected distributions with one-dimensional random variables. In this paper, we show how to construct a diverse family of distributions with known ground-truth mutual information and propose a language-independent benchmarking platform for mutual information estimators. We discuss the general applicability and limitations of classical and neural estimators in settings involving high dimensions, sparse interactions, long-tailed distributions, and high mutual information. Finally, we provide guidelines for practitioners on how to select appropriate estimator adapted to the difficulty of problem considered and issues one needs to consider when applying an estimator to a new data set.
Pawel Czyz, Frederic Grabowski, Julia E. Vogt, Niko Beerenwinkel, Alexander Marx 0001
NeurIPS4
2023 Coherent pathway enrichment estimation by modeling inter-pathway dependencies using regularized regression
abstract
MOTIVATION: Gene set enrichment methods are a common tool to improve the interpretability of gene lists as obtained, for example, from differential gene expression analyses. They are based on computing whether dysregulated genes are located in certain biological pathways more often than expected by chance. Gene set enrichment tools rely on pre-existing pathway databases such as KEGG, Reactome, or the Gene Ontology. These databases are increasing in size and in the number of redundancies between pathways, which complicates the statistical enrichment computation. RESULTS: We address this problem and develop a novel gene set enrichment method, called pareg, which is based on a regularized generalized linear model and directly incorporates dependencies between gene sets related to certain biological functions, for example, due to shared genes, in the enrichment computation. We show that pareg is more robust to noise than competing methods. Additionally, we demonstrate the ability of our method to recover known pathways as well as to suggest novel treatment targets in an exploratory analysis using breast cancer samples from TCGA. AVAILABILITY AND IMPLEMENTATION: pareg is freely available as an R package on Bioconductor (https://bioconductor.org/packages/release/bioc/html/pareg.html) as well as on https://github.com/cbg-ethz/pareg. The GitHub repository also contains the Snakemake workflows needed to reproduce all results presented here.
Kim Philipp Jablonski, Niko Beerenwinkel
Bioinform.2
2023 Predicting tumour content of liquid biopsies from cell-free DNA
abstract
BACKGROUND: Liquid biopsy is a minimally-invasive method of sampling bodily fluids, capable of revealing evidence of cancer. The distribution of cell-free DNA (cfDNA) fragment lengths has been shown to differ between healthy subjects and cancer patients, whereby the distributional shift correlates with the sample's tumour content. These fragmentomic data have not yet been utilised to directly quantify the proportion of tumour-derived cfDNA in a liquid biopsy. RESULTS: We used statistical learning to predict tumour content from Fourier and wavelet transforms of cfDNA length distributions in samples from 118 cancer patients. The model was validated on an independent dilution series of patient plasma. CONCLUSIONS: This proof of concept suggests that our fragmentomic methodology could be useful for predicting tumour content in liquid biopsies.
Mathias Cardner, Francesco Marass, Erika Gedvilaite, Julie L. Yang, Dana W. Y. Tsui, Niko Beerenwinkel
BMC Bioinform.6
2022 Mapping Single-Cell Transcriptomes to Copy Number Evolutionary Trees
Pedro F. Ferreira, Jack Kuipers, Niko Beerenwinkel
RECOMB3
2022 Joint Inference of Repeated Evolutionary Trajectories and Patterns of Clonal Exclusivity or Co-occurrence from Tumor Mutation Trees
Xiang Ge Luo, Jack Kuipers, Niko Beerenwinkel
RECOMB3
2022 Discovering gene regulatory networks of multiple phenotypic groups using dynamic Bayesian networks
abstract
Dynamic Bayesian networks (DBNs) can be used for the discovery of gene regulatory networks (GRNs) from time series gene expression data. Here, we suggest a strategy for learning DBNs from gene expression data by employing a Bayesian approach that is scalable to large networks and is targeted at learning models with high predictive accuracy. Our framework can be used to learn DBNs for multiple groups of samples and highlight differences and similarities in their GRNs. We learn these DBN models based on different structural and parametric assumptions and select the optimal model based on the cross-validated predictive accuracy. We show in simulation studies that our approach is better equipped to prevent overfitting than techniques used in previous studies. We applied the proposed DBN-based approach to two time series transcriptomic datasets from the Gene Expression Omnibus database, each comprising data from distinct phenotypic groups of the same tissue type. In the first case, we used DBNs to characterize responders and non-responders to anti-cancer therapy. In the second case, we compared normal to tumor cells of colorectal tissue. The classification accuracy reached by the DBN-based classifier for both datasets was higher than reported previously. For the colorectal cancer dataset, our analysis suggested that GRNs for cancer and normal tissues have a lot of differences, which are most pronounced in the neighborhoods of oncogenes and known cancer tissue markers. The identified differences in gene networks of cancer and normal cells may be used for the discovery of targeted therapies.
Polina Suter, Jack Kuipers, Niko Beerenwinkel
Briefings Bioinform.3
2022 Identifying cancer pathway dysregulations using differential causal effects
abstract
MOTIVATION: Signaling pathways control cellular behavior. Dysregulated pathways, for example, due to mutations that cause genes and proteins to be expressed abnormally, can lead to diseases, such as cancer. RESULTS: We introduce a novel computational approach, called Differential Causal Effects (dce), which compares normal to cancerous cells using the statistical framework of causality. The method allows to detect individual edges in a signaling pathway that are dysregulated in cancer cells, while accounting for confounding. Hence, technical artifacts have less influence on the results and dce is more likely to detect the true biological signals. We extend the approach to handle unobserved dense confounding, where each latent variable, such as, for example, batch effects or cell cycle states, affects many covariates. We show that dce outperforms competing methods on synthetic datasets and on CRISPR knockout screens. We validate its latent confounding adjustment properties on a GTEx (Genotype-Tissue Expression) dataset. Finally, in an exploratory analysis on breast cancer data from TCGA (The Cancer Genome Atlas), we recover known and discover new genes involved in breast cancer progression. AVAILABILITY AND IMPLEMENTATION: The method dce is freely available as an R package on Bioconductor (https://bioconductor.org/packages/release/bioc/html/dce.html) as well as on https://github.com/cbg-ethz/dce. The GitHub repository also contains the Snakemake workflows needed to reproduce all results presented here. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kim Philipp Jablonski, Martin Pirkl, Domagoj Cevid, Peter Bühlmann, Niko Beerenwinkel
Bioinform.5
2022 Single-cell mutation calling and phylogenetic tree reconstruction with loss and recurrence
abstract
MOTIVATION: Tumours evolve as heterogeneous populations of cells, which may be distinguished by different genomic aberrations. The resulting intra-tumour heterogeneity plays an important role in cancer patient relapse and treatment failure, so that obtaining a clear understanding of each patient's tumour composition and evolutionary history is key for personalized therapies. Single-cell sequencing (SCS) now provides the possibility to resolve tumour heterogeneity at the highest resolution of individual tumour cells, but brings with it challenges related to the particular noise profiles of the sequencing protocols as well as the complexity of the underlying evolutionary process. RESULTS: By modelling the noise processes and allowing mutations to be lost or to reoccur during tumour evolution, we present a method to jointly call mutations in each cell, reconstruct the phylogenetic relationship between cells, and determine the locations of mutational losses and recurrences. Our Bayesian approach allows us to accurately call mutations as well as to quantify our certainty in such predictions. We show the advantages of allowing mutational loss or recurrence with simulated data and present its application to tumour SCS data. AVAILABILITY AND IMPLEMENTATION: SCIΦN is available at https://github.com/cbg-ethz/SCIPhIN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jack Kuipers, Jochen Singer, Niko Beerenwinkel
Bioinform.3
2022 scAmpi - A versatile pipeline for single-cell RNA-seq analysis from basics to clinics
abstract
Single-cell RNA sequencing (scRNA-seq) has emerged as a powerful technique to decipher tissue composition at the single-cell level and to inform on disease mechanisms, tumor heterogeneity, and the state of the immune microenvironment. Although multiple methods for the computational analysis of scRNA-seq data exist, their application in a clinical setting demands standardized and reproducible workflows, targeted to extract, condense, and display the clinically relevant information. To this end, we designed scAmpi (Single Cell Analysis mRNA pipeline), a workflow that facilitates scRNA-seq analysis from raw read processing to informing on sample composition, clinically relevant gene and pathway alterations, and in silico identification of personalized candidate drug treatments. We demonstrate the value of this workflow for clinical decision making in a molecular tumor board as part of a clinical study.
Anne Bertolini, Michael Prummer, Mustafa Anil Tuncel, Ulrike Menzel, María Lourdes Rosano-González, Jack Kuipers, Daniel J. Stekhoven, Niko Beerenwinkel, Franziska Singer
PLoS Comput. Biol.9
2022 Multi-omics subtyping of hepatocellular carcinoma patients using a Bayesian network mixture model
abstract
Comprehensive molecular characterization of cancer subtypes is essential for predicting clinical outcomes and searching for personalized treatments. We present bnClustOmics, a statistical model and computational tool for multi-omics unsupervised clustering, which serves a dual purpose: Clustering patient samples based on a Bayesian network mixture model and learning the networks of omics variables representing these clusters. The discovered networks encode interactions among all omics variables and provide a molecular characterization of each patient subgroup. We conducted simulation studies that demonstrated the advantages of our approach compared to other clustering methods in the case where the generative model is a mixture of Bayesian networks. We applied bnClustOmics to a hepatocellular carcinoma (HCC) dataset comprising genome (mutation and copy number), transcriptome, proteome, and phosphoproteome data. We identified three main HCC subtypes together with molecular characteristics, some of which are associated with survival even when adjusting for the clinical stage. Cluster-specific networks shed light on the links between genotypes and molecular phenotypes of samples within their respective clusters and suggest targets for personalized treatments.
Polina Suter, Eva Dazert, Jack Kuipers, Charlotte K. Y. Ng, Tuyana Boldanova, Michael N. Hall, Markus H. Heim, Niko Beerenwinkel
PLoS Comput. Biol.8
2021 Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus research
abstract
SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected].
Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz
Briefings Bioinform.8
2021 Inferring perturbation profiles of cancer samples
abstract
MOTIVATION: Cancer is one of the most prevalent diseases in the world. Tumors arise due to important genes changing their activity, e.g. when inhibited or over-expressed. But these gene perturbations are difficult to observe directly. Molecular profiles of tumors can provide indirect evidence of gene perturbations. However, inferring perturbation profiles from molecular alterations is challenging due to error-prone molecular measurements and incomplete coverage of all possible molecular causes of gene perturbations. RESULTS: We have developed a novel mathematical method to analyze cancer driver genes and their patient-specific perturbation profiles. We combine genetic aberrations with gene expression data in a causal network derived across patients to infer unobserved perturbations. We show that our method can predict perturbations in simulations, CRISPR perturbation screens and breast cancer samples from The Cancer Genome Atlas. AVAILABILITY AND IMPLEMENTATION: The method is available as the R-package nempi at https://github.com/cbg-ethz/nempi and http://bioconductor.org/packages/nempi. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Martin Pirkl, Niko Beerenwinkel
Bioinform.2
2021 V-pipe: a computational pipeline for assessing viral genetic diversity from high-throughput data
abstract
MOTIVATION: High-throughput sequencing technologies are used increasingly not only in viral genomics research but also in clinical surveillance and diagnostics. These technologies facilitate the assessment of the genetic diversity in intra-host virus populations, which affects transmission, virulence and pathogenesis of viral infections. However, there are two major challenges in analysing viral diversity. First, amplification and sequencing errors confound the identification of true biological variants, and second, the large data volumes represent computational limitations. RESULTS: To support viral high-throughput sequencing studies, we developed V-pipe, a bioinformatics pipeline combining various state-of-the-art statistical models and computational tools for automated end-to-end analyses of raw sequencing reads. V-pipe supports quality control, read mapping and alignment, low-frequency mutation calling, and inference of viral haplotypes. For generating high-quality read alignments, we developed a novel method, called ngshmmalign, based on profile hidden Markov models and tailored to small and highly diverse viral genomes. V-pipe also includes benchmarking functionality providing a standardized environment for comparative evaluations of different pipeline configurations. We demonstrate this capability by assessing the impact of three different read aligners (Bowtie 2, BWA MEM, ngshmmalign) and two different variant callers (LoFreq, ShoRAH) on the performance of calling single-nucleotide variants in intra-host virus populations. V-pipe supports various pipeline configurations and is implemented in a modular fashion to facilitate adaptations to the continuously changing technology landscape. AVAILABILITYAND IMPLEMENTATION: V-pipe is freely available at https://github.com/cbg-ethz/V-pipe. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Susana Posada-Céspedes, David Seifert, Ivan Topolsky, Kim Philipp Jablonski, Karin J. Metzner, Niko Beerenwinkel
Bioinform.6
2021 Statistical tests for intra-tumour clonal co-occurrence and exclusivity
abstract
Tumour progression is an evolutionary process in which different clones evolve over time, leading to intra-tumour heterogeneity. Interactions between clones can affect tumour evolution and hence disease progression and treatment outcome. Intra-tumoural pairs of mutations that are overrepresented in a co-occurring or clonally exclusive fashion over a cohort of patient samples may be suggestive of a synergistic effect between the different clones carrying these mutations. We therefore developed a novel statistical testing framework, called GeneAccord, to identify such gene pairs that are altered in distinct subclones of the same tumour. We analysed our framework for calibration and power. By comparing its performance to baseline methods, we demonstrate that to control type I errors, it is essential to account for the evolutionary dependencies among clones. In applying GeneAccord to the single-cell sequencing of a cohort of 123 acute myeloid leukaemia patients, we find 1 clonally co-occurring and 8 clonally exclusive gene pairs. The clonally exclusive pairs mostly involve genes of the key signalling pathways.
Jack Kuipers, Ariane L. Moore, Katharina Jahn 0001, Peter Schraml, Kiyomi Morita, P. Andrew Futreal, Koichi Takahashi, Christian Beisel, Holger Moch, Niko Beerenwinkel
PLoS Comput. Biol.11
2021 Drug-induced resistance evolution necessitates less aggressive treatment
abstract
Increasing body of experimental evidence suggests that anticancer and antimicrobial therapies may themselves promote the acquisition of drug resistance by increasing mutability. The successful control of evolving populations requires that such biological costs of control are identified, quantified and included to the evolutionarily informed treatment protocol. Here we identify, characterise and exploit a trade-off between decreasing the target population size and generating a surplus of treatment-induced rescue mutations. We show that the probability of cure is maximized at an intermediate dosage, below the drug concentration yielding maximal population decay, suggesting that treatment outcomes may in some cases be substantially improved by less aggressive treatment strategies. We also provide a general analytical relationship that implicitly links growth rate, pharmacodynamics and dose-dependent mutation rate to an optimal control law. Our results highlight the important, but often neglected, role of fundamental eco-evolutionary costs of control. These costs can often lead to situations, where decreasing the cumulative drug dosage may be preferable even when the objective of the treatment is elimination, and not containment. Taken together, our results thus add to the ongoing criticism of the standard practice of administering aggressive, high-dose therapies and motivate further experimental and clinical investigation of the mutagenicity and other hidden collateral costs of therapies.
Teemu Kuosmanen, Johannes Cairns, Robert Noble, Niko Beerenwinkel, Tommi Mononen, Ville Mustonen
PLoS Comput. Biol.4
2021 Comparing mutational pathways to lopinavir resistance in HIV-1 subtypes B versus C
abstract
Although combination antiretroviral therapies seem to be effective at controlling HIV-1 infections regardless of the viral subtype, there is increasing evidence for subtype-specific drug resistance mutations. The order and rates at which resistance mutations accumulate in different subtypes also remain poorly understood. Most of this knowledge is derived from studies of subtype B genotypes, despite not being the most abundant subtype worldwide. Here, we present a methodology for the comparison of mutational networks in different HIV-1 subtypes, based on Hidden Conjunctive Bayesian Networks (H-CBN), a probabilistic model for inferring mutational networks from cross-sectional genotype data. We introduce a Monte Carlo sampling scheme for learning H-CBN models for a larger number of resistance mutations and develop a statistical test to assess differences in the inferred mutational networks between two groups. We apply this method to infer the temporal progression of mutations conferring resistance to the protease inhibitor lopinavir in a large cross-sectional cohort of HIV-1 subtype C genotypes from South Africa, as well as to a data set of subtype B genotypes obtained from the Stanford HIV Drug Resistance Database and the Swiss HIV Cohort Study. We find strong support for different initial mutational events in the protease, namely at residue 46 in subtype B and at residue 82 in subtype C. The inferred mutational networks for subtype B versus C are significantly different sharing only five constraints on the order of accumulating mutations with mutation at residue 54 as the parental event. The results also suggest that mutations can accumulate along various alternative paths within subtypes, as opposed to a unique total temporal ordering. Beyond HIV drug resistance, the statistical methodology is applicable more generally for the comparison of inferred mutational networks between any two groups.
Susana Posada-Céspedes, Gert van Zyl, Hesam Montazeri, Jack Kuipers, Soo-Yon Rhee, Roger D. Kouyos, Huldrych F. Günthard, Niko Beerenwinkel
PLoS Comput. Biol.8
2020 Bayesian Non-parametric Clustering of Single-Cell Mutation Profiles
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
RECOMB6
2020 The Bourque Distances for Mutation Trees of Cancers
abstract
Mutation trees are rooted trees of arbitrary node degree in which each node is labeled with a mutation set. These trees, also referred to as clonal trees, are used in computational oncology to represent the mutational history of tumours. Classical tree metrics such as the popular Robinson - Foulds distance are of limited use for the comparison of mutation trees. One reason is that mutation trees inferred with different methods or for different patients often contain different sets of mutation labels. Here, we generalize the Robinson - Foulds distance into a set of distance metrics called Bourque distances for comparing mutation trees. A connection between the Robinson - Foulds distance and the nearest neighbor interchange distance is also presented.
Katharina Jahn 0001, Niko Beerenwinkel, Louxin Zhang
WABI2
2020 BnpC: Bayesian non-parametric clustering of single-cell mutation profiles
abstract
MOTIVATION: The high resolution of single-cell DNA sequencing (scDNA-seq) offers great potential to resolve intratumor heterogeneity (ITH) by distinguishing clonal populations based on their mutation profiles. However, the increasing size of scDNA-seq datasets and technical limitations, such as high error rates and a large proportion of missing values, complicate this task and limit the applicability of existing methods. RESULTS: Here, we introduce BnpC, a novel non-parametric method to cluster individual cells into clones and infer their genotypes based on their noisy mutation profiles. We benchmarked our method comprehensively against state-of-the-art methods on simulated data using various data sizes, and applied it to three cancer scDNA-seq datasets. On simulated data, BnpC compared favorably against current methods in terms of accuracy, runtime and scalability. Its inferred genotypes were the most accurate, especially on highly heterogeneous data, and it was the only method able to run and produce results on datasets with 5000 cells. On tumor scDNA-seq data, BnpC was able to identify clonal populations missed by the original cluster analysis but supported by Supplementary Experimental Data. With ever growing scDNA-seq datasets, scalable and accurate methods such as BnpC will become increasingly relevant, not only to resolve ITH but also as a preprocessing step to reduce data size. AVAILABILITY AND IMPLEMENTATION: BnpC is freely available under MIT license at https://github.com/cbg-ethz/BnpC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nico Borgsmüller, José Bonet 0001, Francesco Marass, Abel González-Pérez, Núria López-Bigas, Niko Beerenwinkel
Bioinform.6
2020 Host factor prioritization for pan-viral genetic perturbation screens using random intercept models and network propagation
abstract
Genetic perturbation screens using RNA interference (RNAi) have been conducted successfully to identify host factors that are essential for the life cycle of bacteria or viruses. So far, most published studies identified host factors primarily for single pathogens. Furthermore, often only a small subset of genes, e.g., genes encoding kinases, have been targeted. Identification of host factors on a pan-pathogen level, i.e., genes that are crucial for the replication of a diverse group of pathogens has received relatively little attention, despite the fact that such common host factors would be highly relevant, for instance, for devising broad-spectrum anti-pathogenic drugs. Here, we present a novel two-stage procedure for the identification of host factors involved in the replication of different viruses using a combination of random effects models and Markov random walks on a functional interaction network. We first infer candidate genes by jointly analyzing multiple perturbations screens while at the same time adjusting for high variance inherent in these screens. Subsequently the inferred estimates are spread across a network of functional interactions thereby allowing for the analysis of missing genes in the biological studies, smoothing the effect sizes of previously found host factors, and considering a priori pathway information defined over edges of the network. We applied the procedure to RNAi screening data of four different positive-sense single-stranded RNA viruses, Hepatitis C virus, Chikungunya virus, Dengue virus and Severe acute respiratory syndrome coronavirus, and detected novel host factors, including UBC, PLCG1, and DYRK1B, which are predicted to significantly impact the replication cycles of these viruses. We validated the detected host factors experimentally using pharmacological inhibition and an additional siRNA screen and found that some of the predicted host factors indeed influence the replication of these pathogens.
Simon Dirmeier, Christopher Dächert, Martijn van Hemert, Ali Tas, Natacha S. Ogando, Frank J. M. van Kuppeveld, Ralf Bartenschlager, Lars Kaderali, Marco Binder, Niko Beerenwinkel
PLoS Comput. Biol.10
2020 Predicting colorectal cancer risk from adenoma detection via a two-type branching process model
abstract
Despite advances in the modeling and understanding of colorectal cancer development, the dynamics of the progression from benign adenomatous polyp to colorectal carcinoma are still not fully resolved.To take advantage of adenoma size and prevalence data in the National Endoscopic Database of the Clinical Outcomes Research Initiative (CORI) as well as colorectal cancer incidence and size data from the Surveillance Epidemiology and End Results (SEER) database, we construct a two-type branching process model with compartments representing adenoma and carcinoma cells.To perform parameter inference we present a new large-size approximation to the size distribution of the cancer compartment and validate our approach on simulated data.By fitting the model to the CORI and SEER data, we learn biologically relevant parameters, including the transition rate from adenoma to cancer.The inferred parameters allow us to predict the individualized risk of the presence of cancer cells for each screened patient.We provide a web application which allows the user to calculate these individual probabilities at https://ccrc-eth.shinyapps.io/CCRC/.For example, we find a 1 in 100 chance of cancer given the presence of an adenoma between 10 and 20mm size in an average risk patient at age 50.We show that our two-type branching process model recapitulates the early growth dynamics of colon adenomas and cancers and can recover epidemiological trends such as adenoma prevalence and cancer incidence while remaining mathematically and computationally tractable. Author summaryColorectal cancer is a major public health burden.The development of colorectal cancer starts with the mutational initiation of non-cancerous growths in the form of benign adenomatous polyps.These adenomas grow over time with the potential to develop into carcinomas.Many mathematical and simulation-based models have been used to gain insight into this process.We aimed to understand rates of adenoma growth and transition
Brian Lang, Jack Kuipers, Benjamin Misselwitz, Niko Beerenwinkel
PLoS Comput. Biol.4
2020 A mathematical model of the metastatic bottleneck predicts patient outcome and response to cancer treatment
abstract
Metastases are the main reason for cancer-related deaths. Initiation of metastases, where newly seeded tumor cells expand into colonies, presents a tremendous bottleneck to metastasis formation. Despite its importance, a quantitative description of metastasis initiation and its clinical implications is lacking. Here, we set theoretical grounds for the metastatic bottleneck with a simple stochastic model. The model assumes that the proliferation-to-death rate ratio for the initiating metastatic cells increases when they are surrounded by more of their kind. For a total of 159,191 patients across 13 cancer types, we found that a single cell has an extremely low median probability of successful seeding of the order of 10-8. With increasing colony size, a sharp transition from very unlikely to very likely successful metastasis initiation occurs. The median metastatic bottleneck, defined as the critical colony size that marks this transition, was between 10 and 21 cells. We derived the probability of metastasis occurrence and patient outcome based on primary tumor size at diagnosis and tumor type. The model predicts that the efficacy of patient treatment depends on the primary tumor size but even more so on the severity of the metastatic bottleneck, which is estimated to largely vary between patients. We find that medical interventions aiming at tightening the bottleneck, such as immunotherapy, can be much more efficient than therapies that decrease overall tumor burden, such as chemotherapy.
Ewa Szczurek, Tyll Krüger, Barbara Klink, Niko Beerenwinkel
PLoS Comput. Biol.4
2019 Bioinformatics for precision oncology
abstract
Molecular profiling of tumor biopsies plays an increasingly important role not only in cancer research, but also in the clinical management of cancer patients. Multi-omics approaches hold the promise of improving diagnostics, prognostics and personalized treatment. To deliver on this promise of precision oncology, appropriate bioinformatics methods for managing, integrating and analyzing large and complex data are necessary. Here, we discuss the specific requirements of bioinformatics methods and software that arise in the setting of clinical oncology, owing to a stricter regulatory environment and the need for rapid, highly reproducible and robust procedures. We describe the workflow of a molecular tumor board and the specific bioinformatics support that it requires, from the primary analysis of raw molecular profiling data to the automatic generation of a clinical report and its delivery to decision-making clinical oncologists. Such workflows have to various degrees been implemented in many clinical trials, as well as in molecular tumor boards at specialized cancer centers and university hospitals worldwide. We review these and more recent efforts to include other high-dimensional multi-omics patient profiles into the tumor board, as well as the state of clinical decision support software to translate molecular findings into treatment recommendations.
Jochen Singer, Anja Irmisch, Hans-Joachim Ruscheweyh, Franziska Singer, Nora C. Toussaint, Mitchell P. Levesque, Daniel J. Stekhoven, Niko Beerenwinkel
Briefings Bioinform.8
2019 Inferring signalling dynamics by integrating interventional with observational data
abstract
MOTIVATION: In order to infer a cell signalling network, we generally need interventional data from perturbation experiments. If the perturbation experiments are time-resolved, then signal progression through the network can be inferred. However, such designs are infeasible for large signalling networks, where it is more common to have steady-state perturbation data on the one hand, and a non-interventional time series on the other. Such was the design in a recent experiment investigating the coordination of epithelial-mesenchymal transition (EMT) in murine mammary gland cells. We aimed to infer the underlying signalling network of transcription factors and microRNAs coordinating EMT, as well as the signal progression during EMT. RESULTS: In the context of nested effects models, we developed a method for integrating perturbation data with a non-interventional time series. We applied the model to RNA sequencing data obtained from an EMT experiment. Part of the network inferred from RNA interference was validated experimentally using luciferase reporter assays. Our model extension is formulated as an integer linear programme, which can be solved efficiently using heuristic algorithms. This extension allowed us to infer the signal progression through the network during an EMT time course, and thereby assess when each regulator is necessary for EMT to advance. AVAILABILITY AND IMPLEMENTATION: R package at https://github.com/cbg-ethz/timeseriesNEM. The RNA sequencing data and microscopy images can be explored through a Shiny app at https://emt.bsse.ethz.ch. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mathias Cardner, Nathalie Meyer-Schaller, Gerhard Christofori, Niko Beerenwinkel
Bioinform.4
2019 Estimating the predictability of cancer evolution
abstract
MOTIVATION: How predictable is the evolution of cancer? This fundamental question is of immense relevance for the diagnosis, prognosis and treatment of cancer. Evolutionary biologists have approached the question of predictability based on the underlying fitness landscape. However, empirical fitness landscapes of tumor cells are impossible to determine in vivo. Thus, in order to quantify the predictability of cancer evolution, alternative approaches are required that circumvent the need for fitness landscapes. RESULTS: We developed a computational method based on conjunctive Bayesian networks (CBNs) to quantify the predictability of cancer evolution directly from mutational data, without the need for measuring or estimating fitness. Using simulated data derived from >200 different fitness landscapes, we show that our CBN-based notion of evolutionary predictability strongly correlates with the classical notion of predictability based on fitness landscapes under the strong selection weak mutation assumption. The statistical framework enables robust and scalable quantification of evolutionary predictability. We applied our approach to driver mutation data from the TCGA and the MSK-IMPACT clinical cohorts to systematically compare the predictability of 15 different cancer types. We found that cancer evolution is remarkably predictable as only a small fraction of evolutionary trajectories are feasible during cancer progression. AVAILABILITY AND IMPLEMENTATION: https://github.com/cbg-ethz/predictability\_of\_cancer\_evolution. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sayed-Rzgar Hosseini, Ramón Díaz-Uriarte, Florian Markowetz, Niko Beerenwinkel
Bioinform.4
2019 PyBDA: a command line tool for automated analysis of big biological data sets
abstract
BACKGROUND: Analysing large and high-dimensional biological data sets poses significant computational difficulties for bioinformaticians due to lack of accessible tools that scale to hundreds of millions of data points. RESULTS: We developed a novel machine learning command line tool called PyBDA for automated, distributed analysis of big biological data sets. By using Apache Spark in the backend, PyBDA scales to data sets beyond the size of current applications. It uses Snakemake in order to automatically schedule jobs to a high-performance computing cluster. We demonstrate the utility of the software by analyzing image-based RNA interference data of 150 million single cells. CONCLUSION: PyBDA allows automated, easy-to-use data analysis using common statistical methods and machine learning algorithms. It can be used with simple command line calls entirely making it accessible to a broad user base. PyBDA is available at https://pybda.rtfd.io.
Simon Dirmeier, Mario Emmenlauer, Christoph Dehio, Niko Beerenwinkel
BMC Bioinform.4
2018 Integrative Inference of Subclonal Tumour Evolution from Single-Cell and Bulk Sequencing Data
Salem Malikic, Katharina Jahn 0001, Jack Kuipers, Süleyman Cenk Sahinalp, Niko Beerenwinkel
RECOMB5
2018 ModulOmics: Integrating Multi-Omics Data to Identify Cancer Driver Modules
Dana Silverbush, Simona Cristea, Gali Yanovich, Tamar Geiger, Niko Beerenwinkel, Roded Sharan
RECOMB5
2018 SCIΦ: Single-Cell Mutation Identification via Phylogenetic Inference
Jochen Singer, Jack Kuipers, Katharina Jahn 0001, Niko Beerenwinkel
RECOMB4
2018 Network-based integration of multi-omics data for prioritizing cancer genes
abstract
Motivation: Several molecular events are known to be cancer-related, including genomic aberrations, hypermethylation of gene promoter regions and differential expression of microRNAs. These aberration events are very heterogeneous across tumors and it is poorly understood how they affect the molecular makeup of the cell, including the transcriptome and proteome. Protein interaction networks can help decode the functional relationship between aberration events and changes in gene and protein expression. Results: We developed NetICS (Network-based Integration of Multi-omics Data), a new graph diffusion-based method for prioritizing cancer genes by integrating diverse molecular data types on a directed functional interaction network. NetICS prioritizes genes by their mediator effect, defined as the proximity of the gene to upstream aberration events and to downstream differentially expressed genes and proteins in an interaction network. Genes are prioritized for individual samples separately and integrated using a robust rank aggregation technique. NetICS provides a comprehensive computational framework that can aid in explaining the heterogeneity of aberration events by their functional convergence to common differentially expressed genes and proteins. We demonstrate NetICS' competitive performance in predicting known cancer genes and in generating robust gene lists using TCGA data from five cancer types. Availability and implementation: NetICS is available at https://github.com/cbg-ethz/netics. Supplementary information: Supplementary data are available at Bioinformatics online.
Christos Dimitrakopoulos, Sravanth Kumar Hindupur, Luca Häfliger, Jonas Behr, Hesam Montazeri, Michael N. Hall, Niko Beerenwinkel
Bioinform.7
2018 Single cell network analysis with a mixture of Nested Effects Models
abstract
Motivation: New technologies allow for the elaborate measurement of different traits of single cells under genetic perturbations. These interventional data promise to elucidate intra-cellular networks in unprecedented detail and further help to improve treatment of diseases like cancer. However, cell populations can be very heterogeneous. Results: We developed a mixture of Nested Effects Models (M&NEM) for single-cell data to simultaneously identify different cellular subpopulations and their corresponding causal networks to explain the heterogeneity in a cell population. For inference, we assign each cell to a network with a certain probability and iteratively update the optimal networks and cell probabilities in an Expectation Maximization scheme. We validate our method in the controlled setting of a simulation study and apply it to three data sets of pooled CRISPR screens generated previously by two novel experimental techniques, namely Crop-Seq and Perturb-Seq. Availability and implementation: The mixture Nested Effects Model (M&NEM) is available as the R-package mnem at https://github.com/cbg-ethz/mnem/. Supplementary information: Supplementary data are available at Bioinformatics online.
Martin Pirkl, Niko Beerenwinkel
Bioinform.2
2018 NGS-pipe: a flexible, easily extendable and highly configurable framework for NGS analysis
abstract
Motivation: Next-generation sequencing is now an established method in genomics, and massive amounts of sequencing data are being generated on a regular basis. Analysis of the sequencing data is typically performed by lab-specific in-house solutions, but the agreement of results from different facilities is often small. General standards for quality control, reproducibility and documentation are missing. Results: We developed NGS-pipe, a flexible, transparent and easy-to-use framework for the design of pipelines to analyze whole-exome, whole-genome and transcriptome sequencing data. NGS-pipe facilitates the harmonization of genomic data analysis by supporting quality control, documentation, reproducibility, parallelization and easy adaptation to other NGS experiments. Availability and implementation: https://github.com/cbg-ethz/NGS-pipe. Contact: [email protected].
Jochen Singer, Hans-Joachim Ruscheweyh, Ariane L. Hofmann, Thomas Thurnherr, Franziska Singer, Nora C. Toussaint, Charlotte K. Y. Ng, Salvatore Piscuoglio, Christian Beisel, Gerhard Christofori, Reinhard Dummer, Michael N. Hall, Wilhelm Krek, Mitchell P. Levesque, Markus G. Manz, Holger Moch, Andreas Papassotiropoulos, Daniel J. Stekhoven, Peter Wild, Thomas Wüst, Bernd Rinn, Niko Beerenwinkel
Bioinform.22
2018 Improved pathway reconstruction from RNA interference screens by exploiting off-target effects
abstract
Motivation: Pathway reconstruction has proven to be an indispensable tool for analyzing the molecular mechanisms of signal transduction underlying cell function. Nested effects models (NEMs) are a class of probabilistic graphical models designed to reconstruct signalling pathways from high-dimensional observations resulting from perturbation experiments, such as RNA interference (RNAi). NEMs assume that the short interfering RNAs (siRNAs) designed to knockdown specific genes are always on-target. However, it has been shown that most siRNAs exhibit strong off-target effects, which further confound the data, resulting in unreliable reconstruction of networks by NEMs. Results: Here, we present an extension of NEMs called probabilistic combinatorial nested effects models (pc-NEMs), which capitalize on the ancillary siRNA off-target effects for network reconstruction from combinatorial gene knockdown data. Our model employs an adaptive simulated annealing search algorithm for simultaneous inference of network structure and error rates inherent to the data. Evaluation of pc-NEMs on simulated data with varying number of phenotypic effects and noise levels as well as real data demonstrates improved reconstruction compared to classical NEMs. Application to Bartonella henselae infection RNAi screening data yielded an eight node network largely in agreement with previous works, and revealed novel binary interactions of direct impact between established components. Availability and implementation: The software used for the analysis is freely available as an R package at https://github.com/cbg-ethz/pcNEM.git. Supplementary information: Supplementary data are available at Bioinformatics online.
Sumana Srivatsa, Jack Kuipers, Fabian Schmich, Simone Eicher, Mario Emmenlauer, Christoph Dehio, Niko Beerenwinkel
Bioinform.7
2018 Ensemble outlier detection and gene selection in triple-negative breast cancer data
abstract
BACKGROUND: Learning accurate models from 'omics data is bringing many challenges due to their inherent high-dimensionality, e.g. the number of gene expression variables, and comparatively lower sample sizes, which leads to ill-posed inverse problems. Furthermore, the presence of outliers, either experimental errors or interesting abnormal clinical cases, may severely hamper a correct classification of patients and the identification of reliable biomarkers for a particular disease. We propose to address this problem through an ensemble classification setting based on distinct feature selection and modeling strategies, including logistic regression with elastic net regularization, Sparse Partial Least Squares - Discriminant Analysis (SPLS-DA) and Sparse Generalized PLS (SGPLS), coupled with an evaluation of the individuals' outlierness based on the Cook's distance. The consensus is achieved with the Rank Product statistics corrected for multiple testing, which gives a final list of sorted observations by their outlierness level. RESULTS: We applied this strategy for the classification of Triple-Negative Breast Cancer (TNBC) RNA-Seq and clinical data from the Cancer Genome Atlas (TCGA). The detected 24 outliers were identified as putative mislabeled samples, corresponding to individuals with discrepant clinical labels for the HER2 receptor, but also individuals with abnormal expression values of ER, PR and HER2, contradictory with the corresponding clinical labels, which may invalidate the initial TNBC label. Moreover, the model consensus approach leads to the selection of a set of genes that may be linked to the disease. These results are robust to a resampling approach, either by selecting a subset of patients or a subset of genes, with a significant overlap of the outlier patients identified. CONCLUSIONS: The proposed ensemble outlier detection approach constitutes a robust procedure to identify abnormal cases and consensus covariates, which may improve biomarker selection for precision medicine applications. The method can also be easily extended to other regression models and datasets.
Marta B. Lopes, André Veríssimo, Eunice Carrasquinha, Sandra Casimiro, Niko Beerenwinkel, Susana Vinga
BMC Bioinform.5
2017 ISMB/ECCB 2017 proceedings
abstract
The 25th Annual Conference Intelligent Systems for Molecular Biology (ISMB), held jointly with the 16th annual European Conference on Computational Biology (ECCB), took place in Prague, Czech Republic, July 21–25, 2017 (http://www.iscb.org/ismbeccb2017). This special issue of Bioinformatics serves as the proceedings for this official conference of the International Society for Computational Biology (ISCB, http://www.iscb.org/). The year 2017 was inaugural for a new program format introducing the ISCB Communities of Special Interest (COSIs) as scientific themes within the conference. The 12 COSIs presenting at ISMB/ECCB 2017 (Table 1) had previously organized the pre-conference Special Interest Group meetings. This year, the program incorporated COSI sessions including proceedings, highlights (previously published research), and late-breaking research talks in addition to special sessions, workshop, posters and other scientific talks submitted directly to ISMB/ECCB. Thematic areas of ISMB/ECCB 2017. Areas have been defined as Communities of Special Interest (COSIs), except ‘Other Topics’ Thematic areas of ISMB/ECCB 2017. Areas have been defined as Communities of Special Interest (COSIs), except ‘Other Topics’ ISMB/ECCB is the leading international forum in the field for presenting new methods and research results and for facilitating discussions at all levels of expertise. The 45 papers in this volume were selected from 279 original submissions, resulting in an acceptance rate of 16%. Submitted papers accumulated 3.2 reviews on average. The review process included 208 expert reviewers (with some reviewers working for more than one COSI) led by 42 Area Chairs (Table 1) to produce 899 reviews. The Area Chairs were appointed by the individual COSIs, and included a mix of experienced individuals intimately familiar with the field as well as with the review process. Manuscripts submitted to the proceedings track were assigned to COSIs based on authors’ suggestions and on the judgement of Area and Proceedings Chairs. An additional track was created to house all submissions unrelated to any of the existing COSI topics. The submissions in this ‘Other Topics’ area mostly addressed questions in microbiome analysis and evolutionary modeling. With the COSI-based structure, proceedings areas have become more community-driven and dynamic. We believe that going forward in this manner new emerging scientific themes can be detected quicker and represented better. After initial screening, submissions were sent out to expert reviewers selected by the Area Chairs from the program committee. Each paper and its reviews were then discussed online among reviewers and Area Chairs. Finally, Area and Proceedings Chairs discussed the ranked papers to arrive at a subset of 46 papers, which were conditionally accepted. In continuation of the two-tier review process employed at ISMB 2016, Area Chairs, Proceedings Chairs and sometimes the corresponding reviewers inspected the revised papers again. All but one paper were re-submitted. These 45 revised papers were judged to have satisfactorily addressed the reviewers’ comments and were finally accepted into the proceedings. We would like to thank all authors of the 279 submissions for submitting their work. The excellent scientific quality of these papers is pivotal for the impact and the success of the conference. We believe that we have identified 45 outstanding pieces of bioinformatics research for the proceedings track. We note, however, that the review process was subject to some re-organization largely due to the new COSI-based conference format. It is, therefore, possible that excellent papers have not been fully recognized as such. Nevertheless, we know that submissions were evaluated fairly and diligently and hope that all authors have received useful feedback. Like the conference itself, the review and selection process of the proceedings papers is a community effort. We would like to express our deep gratitude to all Area Chairs, members of the program committee, and external subreviewers for their outstanding efforts and rigorous focus on scientific quality. We would like to express deep gratitude to Steven Leard for all help with the review process—without his organizational support our work would not be possible. We would also like to thank the team at Oxford University Press for consistently handling this ISMB/ECCB special proceedings issue. Last, but certainly not least, we thank all members of the ISMB Steering Committee for their expert advice and support. Conflict of Interest: none declared.
Niko Beerenwinkel, Yana Bromberg
Bioinform.1
2017 Detailed simulation of cancer exome sequencing data reveals differences and common limitations of variant callers
abstract
BACKGROUND: Next-generation sequencing of matched tumor and normal biopsy pairs has become a technology of paramount importance for precision cancer treatment. Sequencing costs have dropped tremendously, allowing the sequencing of the whole exome of tumors for just a fraction of the total treatment costs. However, clinicians and scientists cannot take full advantage of the generated data because the accuracy of analysis pipelines is limited. This particularly concerns the reliable identification of subclonal mutations in a cancer tissue sample with very low frequencies, which may be clinically relevant. RESULTS: Using simulations based on kidney tumor data, we compared the performance of nine state-of-the-art variant callers, namely deepSNV, GATK HaplotypeCaller, GATK UnifiedGenotyper, JointSNVMix2, MuTect, SAMtools, SiNVICT, SomaticSniper, and VarScan2. The comparison was done as a function of variant allele frequencies and coverage. Our analysis revealed that deepSNV and JointSNVMix2 perform very well, especially in the low-frequency range. We attributed false positive and false negative calls of the nine tools to specific error sources and assigned them to processing steps of the pipeline. All of these errors can be expected to occur in real data sets. We found that modifying certain steps of the pipeline or parameters of the tools can lead to substantial improvements in performance. Furthermore, a novel integration strategy that combines the ranks of the variants yielded the best performance. More precisely, the rank-combination of deepSNV, JointSNVMix2, MuTect, SiNVICT and VarScan2 reached a sensitivity of 78% when fixing the precision at 90%, and outperformed all individual tools, where the maximum sensitivity was 71% with the same precision. CONCLUSIONS: The choice of well-performing tools for alignment and variant calling is crucial for the correct interpretation of exome sequencing data obtained from mixed samples, and common pipelines are suboptimal. We were able to relate observed substantial differences in performance to the underlying statistical models of the tools, and to pinpoint the error sources of false positive and false negative calls. These findings might inspire new software developments that improve exome sequencing pipelines and further the field of precision cancer treatment.
Ariane L. Hofmann, Jonas Behr, Jochen Singer, Jack Kuipers, Christian Beisel, Peter Schraml, Holger Moch, Niko Beerenwinkel
BMC Bioinform.8
2017 Limb-Enhancer Genie: An accessible resource of accurate enhancer predictions in the developing limb
abstract
Epigenomic mapping of enhancer-associated chromatin modifications facilitates the genome-wide discovery of tissue-specific enhancers in vivo. However, reliance on single chromatin marks leads to high rates of false-positive predictions. More sophisticated, integrative methods have been described, but commonly suffer from limited accessibility to the resulting predictions and reduced biological interpretability. Here we present the Limb-Enhancer Genie (LEG), a collection of highly accurate, genome-wide predictions of enhancers in the developing limb, available through a user-friendly online interface. We predict limb enhancers using a combination of >50 published limb-specific datasets and clusters of evolutionarily conserved transcription factor binding sites, taking advantage of the patterns observed at previously in vivo validated elements. By combining different statistical models, our approach outperforms current state-of-the-art methods and provides interpretable measures of feature importance. Our results indicate that including a previously unappreciated score that quantifies tissue-specific nuclease accessibility significantly improves prediction performance. We demonstrate the utility of our approach through in vivo validation of newly predicted elements. Moreover, we describe general features that can guide the type of datasets to include when predicting tissue-specific enhancers genome-wide, while providing an accessible resource to the general biological community and facilitating the functional interpretation of genetic studies of limb malformations.
Remo Monti, Iros Barozzi, Marco Osterwalder, Elizabeth Lee, Momoe Kato, Tyler H. Garvin, Ingrid Plajzer-Frick, Catherine S. Pickle, Jennifer A. Akiyama, Veena Afzal, Niko Beerenwinkel, Diane E. Dickel, Axel Visel, Len A. Pennacchio
PLoS Comput. Biol.11
2017 Inferring modulators of genetic interactions with epistatic nested effects models
abstract
Maps of genetic interactions can dissect functional redundancies in cellular networks. Gene expression profiles as high-dimensional molecular readouts of combinatorial perturbations provide a detailed view of genetic interactions, but can be hard to interpret if different gene sets respond in different ways (called mixed epistasis). Here we test the hypothesis that mixed epistasis between a gene pair can be explained by the action of a third gene that modulates the interaction. We have extended the framework of Nested Effects Models (NEMs), a type of graphical model specifically tailored to analyze high-dimensional gene perturbation data, to incorporate logical functions that describe interactions between regulators on downstream genes and proteins. We benchmark our approach in the controlled setting of a simulation study and show high accuracy in inferring the correct model. In an application to data from deletion mutants of kinases and phosphatases in S. cerevisiae we show that epistatic NEMs can point to modulators of genetic interactions. Our approach is implemented in the R-package 'epiNEM' available from https://github.com/cbg-ethz/epiNEM and https://bioconductor.org/packages/epiNEM/.
Martin Pirkl, Madeline Diekmann, Marlies van der Wees, Niko Beerenwinkel, Holger Fröhlich, Florian Markowetz
PLoS Comput. Biol.4
2016 pathTiMEx: Joint Inference of Mutually Exclusive Cancer Pathways and Their Dependencies in Tumor Progression
Simona Cristea, Jack Kuipers, Niko Beerenwinkel
RECOMB3
2016 Tree Inference for Single-Cell Data
Katharina Jahn 0001, Jack Kuipers, Niko Beerenwinkel
RECOMB3
2016 TiMEx: a waiting time model for mutually exclusive cancer alterations
abstract
MOTIVATION: Despite recent technological advances in genomic sciences, our understanding of cancer progression and its driving genetic alterations remains incomplete. RESULTS: We introduce TiMEx, a generative probabilistic model for detecting patterns of various degrees of mutual exclusivity across genetic alterations, which can indicate pathways involved in cancer progression. TiMEx explicitly accounts for the temporal interplay between the waiting times to alterations and the observation time. In simulation studies, we show that our model outperforms previous methods for detecting mutual exclusivity. On large-scale biological datasets, TiMEx identifies gene groups with strong functional biological relevance, while also proposing new candidates for biological validation. TiMEx possesses several advantages over previous methods, including a novel generative probabilistic model of tumorigenesis, direct estimation of the probability of mutual exclusivity interaction, computational efficiency and high sensitivity in detecting gene groups involving low-frequency alterations. AVAILABILITY AND IMPLEMENTATION: TiMEx is available as a Bioconductor R package at www.bsse.ethz.ch/cbg/software/TiMEx CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Simona Constantinescu, Ewa Szczurek, Pejman Mohammadi 0001, Jörg Rahnenführer, Niko Beerenwinkel
Bioinform.5
2016 BMix: probabilistic modeling of occurring substitutions in PAR-CLIP data
abstract
MOTIVATION: Photoactivatable ribonucleoside-enhanced cross-linking and immunoprecipitation (PAR-CLIP) is an experimental method based on next-generation sequencing for identifying the RNA interaction sites of a given protein. The method deliberately inserts T-to-C substitutions at the RNA-protein interaction sites, which provides a second layer of evidence compared with other CLIP methods. However, the experiment includes several sources of noise which cause both low-frequency errors and spurious high-frequency alterations. Therefore, rigorous statistical analysis is required in order to separate true T-to-C base changes, following cross-linking, from noise. So far, most of the existing PAR-CLIP data analysis methods focus on discarding the low-frequency errors and rely on high-frequency substitutions to report binding sites, not taking into account the possibility of high-frequency false positive substitutions. RESULTS: Here, we introduce BMix, a new probabilistic method which explicitly accounts for the sources of noise in PAR-CLIP data and distinguishes cross-link induced T-to-C substitutions from low and high-frequency erroneous alterations. We demonstrate the superior speed and accuracy of our method compared with existing approaches on both simulated and real, publicly available human datasets. AVAILABILITY AND IMPLEMENTATION: The model is freely accessible within the BMix toolbox at www.cbg.bsse.ethz.ch/software/BMix, available for Matlab and R. SUPPLEMENTARY INFORMATION: Supplementary data is available at Bioinformatics online. CONTACT: [email protected].
Monica Golumbeanu, Pejman Mohammadi 0001, Niko Beerenwinkel
Bioinform.3
2016 Large-scale inference of conjunctive Bayesian networks
abstract
UNLABELLED: The continuous time conjunctive Bayesian network (CT-CBN) is a graphical model for analyzing the waiting time process of the accumulation of genetic changes (mutations). CT-CBN models have been successfully used in several biological applications such as HIV drug resistance development and genetic progression of cancer. However, current approaches for parameter estimation and network structure learning of CBNs can only deal with a small number of mutations (<20). Here, we address this limitation by presenting an efficient and accurate approximate inference algorithm using a Monte Carlo expectation-maximization algorithm based on importance sampling. The new method can now be used for a large number of mutations, up to one thousand, an increase by two orders of magnitude. In simulation studies, we present the accuracy as well as the running time efficiency of the new inference method and compare it with a MLE method, expectation-maximization, and discrete time CBN model, i.e. a first-order approximation of the CT-CBN model. We also study the application of the new model on HIV drug resistance datasets for the combination therapy with zidovudine plus lamivudine (AZT + 3TC) as well as under no treatment, both extracted from the Swiss HIV Cohort Study database. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented as an R package available at https://github.com/cbg-ethz/MC-CBN CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hesam Montazeri, Jack Kuipers, Roger D. Kouyos, Jürg Böni, Sabine Yerly, Thomas Klimkait, Vincent Aubert, Huldrych F. Günthard, Niko Beerenwinkel
Bioinform.9
2016 Linear effects models of signaling pathways from combinatorial perturbation data
abstract
MOTIVATION: Perturbations constitute the central means to study signaling pathways. Interrupting components of the pathway and analyzing observed effects of those interruptions can give insight into unknown connections within the signaling pathway itself, as well as the link from the pathway to the effects. Different pathway components may have different individual contributions to the measured perturbation effects, such as gene expression changes. Those effects will be observed in combination when the pathway components are perturbed. Extant approaches focus either on the reconstruction of pathway structure or on resolving how the pathway components control the downstream effects. RESULTS: Here, we propose a linear effects model, which can be applied to solve both these problems from combinatorial perturbation data. We use simulated data to demonstrate the accuracy of learning the pathway structure as well as estimation of the individual contributions of pathway components to the perturbation effects. The practical utility of our approach is illustrated by an application to perturbations of the mitogen-activated protein kinase pathway in Saccharomyces cerevisiaeAvailability and Implementation: lem is available as a R package at http://www.mimuw.edu.pl/∼szczurek/lem CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ewa Szczurek, Niko Beerenwinkel
Bioinform.2
2016 Computational Cancer Biology: An Evolutionary Perspective
abstract
Cancer is a leading cause of death worldwide and represents one of the biggest biomedical research challenges of our time.Tumor progression is caused by somatic evolution of cell populations.Cancer cells expand because of the accumulation of selectively advantageous mutations, and expanding clones give rise to new cell subpopulations with increasingly higher somatic fitness (Fig 1).In the 1970s, Nowell and others established this somatic evolutionary view of cancer [1].Today, computational biologists have the opportunity to take advantage of large-scale molecular profiling data in order to carve out the principles of tumor evolution and to elucidate how it manifests across cancer types.Analogous to other evolutionary studies, mathematical modeling will be key to the success of understanding the somatic evolution of cancer [2].In general, cancer research involves a range of clinical, epidemiological, and molecular approaches, as well as mathematical and computational modeling.An early and very successful example of mathematical modeling was the work of Nordling [3] and of Armitage and Doll [4].In the 1950s, long before cancer genome data was available, they analyzed cancer incidence data and postulated, based on the observed age-incidence curves, that cancer is a multistep process.In search of these rate-limiting events, cancer progression was then linked to the accumulation of genomic alterations.Since then, the evolutionary perspective on cancer has proven useful in many instances, and the mathematical theory of cancer evolution has been developed much further.However, little clinical benefit could be gained from this approach so far.Much of evolutionary modeling in general, and of cancer in particular, has remained conceptual or qualitative, either because of strong simplifications in the interest of mathematical tractability or lack of informative data.Next-generation sequencing (NGS) technologies and their various applications have changed this situation fundamentally [5].Today, cancer cells can be analyzed in great detail at the molecular level, and tumor cell populations can be sampled extensively.Driven by this technological revolution, large numbers of high-dimensional molecular profiles of tumors, and even of individual cancer cells, are collected by cancer genome consortia, as well as by many individual labs.Large catalogs of cancer genomes, epigenomes, transcriptomes, proteomes, and other molecular profiles are generated to assess variation among tumors from different patients (intertumor heterogeneity) as well as among individual cells of single tumors (intratumor heterogeneity).These data hold the promise not only of new cancer biology discoveries but also of progress in cancer diagnostics and treatment.Analyzing these complex data and interpreting them in the context of ongoing somatic evolution, disease progression, and treatment response is a major challenge, and the prospects to
Niko Beerenwinkel, Chris D. Greenman, Jens Lagergren
PLoS Comput. Biol.1
2015 Finding Dense Subgraphs in Relational Graphs
Vinay Jethava, Niko Beerenwinkel
ECML/PKDD (2)2
2015 ISMB/ECCB 2015
abstract
This special issue of Bioinformatics serves as the proceedings of the joint 23rd annual meeting of Intelligent Systems for Molecular Biology (ISMB) and 14th European Conference on Computational Biology (ECCB), which took place in Dublin, Ireland, July 10–14, 2015 (http://www.iscb.org/ismbeccb2015). ISMB/ECCB 2015, the official conference of the International Society for Computational Biology (ISCB, http://www.iscb.org/), was accompanied by nine Special Interest Group meetings of 1 or 2 days each, and two satellite meetings. Since its inception, ISMB/ECCB has been the largest international conference in computational biology and bioinformatics. It is the leading forum in the field for presenting new research results, disseminating methods and techniques, and facilitating discussions among leading researchers, practitioners, and students in the field. The 42 papers in this volume were selected from 241 original submissions divided into 13 research areas, collectively led by 25 Area Chairs. For each area, the Area Chairs selected an expert program committee for their subdiscipline and oversaw the reviewing process for that area. By design, the Area Chairs included a mix of experienced individuals reappointed from previous years and experts newly recruited to ensure broad technical expertise and to promote inclusivity of various elements of the research community. In total, the review process involved the 25 Area Chairs, 378 program committee members, and an additional 175 external reviewers recruited as sub-reviewers by program committee members. Table 1 provides a summary of the areas, area chairs and a review summary by area. The conference used a two-tier review system—a continuation and refinement of a process that begun with ISMB/ECCB 2013 in an effort to better ensure thorough and fair reviewing. Under the revised process, each of the 241 submissions was first reviewed by at least three expert referees, with a subset receiving between four and six reviews, as needed. Consensus on each paper was reached through online discussion among reviewers and Area Chairs. Among the 241 submissions, 27 were conditionally accepted for publication directly from the first round review. A subset of 29 papers was viewed as potentially publishable subject to revision and re-review of the manuscripts. Of the 27 papers that were resubmitted, 15 were judged to have addressed the concerns of the reviewers and were accepted for the conference proceedings, resulting in a total of 42 acceptances and an overall acceptance rate of 42/241 = 17.4%. We believe that this two-tier system, which is more reflective of typical multi-round journal review procedures, provided a means of ensuring that only the highest quality original work was accepted within the tight timing constraints imposed by the conference scheduling. We thank all authors for submitting their work. These proceedings would simply not be possible without the scientific ingenuity of the contributors of all the papers. We recognize that the process is not perfect, and some outstanding work might have been rejected despite our best efforts. Nonetheless, we are hopeful that all authors received helpful feedback on their work and that most believed their submissions were judged fairly and diligently. In total, the two-tier review process involved 896 individual reviews. We are immensely grateful to the Area Chairs, the members of the program committee and the external subreviewers for their outstanding efforts in conducting a thorough review process in just 3 months. Their contribution is at the core of the scientific quality of the conference. We also thank Steven Leard for his continuing support with the review process; the team at Oxford University Press for preparing this special proceedings volume and the Conference Chairs, Janet Kelso and Alex Bateman, the Theme Chairs and other members of the ISMB Steering Committee for their advice and supervision. We are also grateful to Russell Schwartz, Proceedings Chair of ISMB 2014, for sharing his experience and various helpful documents on the review process. ISMB/ECCB 2015 review summary by area. ISMB/ECCB 2015 review summary by area.
Yves Moreau, Niko Beerenwinkel
Bioinform.2
2015 Covering Pairs in Directed Acyclic Graphs
abstract
The Minimum Path Cover (MinPC) problem on directed acyclic graphs (DAGs) is a classical problem in graph theory that provides a clear and simple mathematical formulation for several applications in computational biology. In this paper, we study the computational complexity of three constrained variants of MinPC motivated by the recent introduction of Next-Generation Sequencing technologies. The first variant (MinRPC), given a DAG and a set of pairs of vertices, asks for a minimum-cardinality set of (not necessarily disjoint) paths such that both vertices of each pair belong to the same path. For this problem, we establish a sharp tractability borderline depending on the ‘overlapping degree’ of the instance, a natural parameter in some applications of the problem. The second variant we consider (MinPCRP), given a DAG and a set of pairs of vertices, asks for a minimum-cardinality set of (not necessarily disjoint) paths ‘covering’ all the vertices of the graph and such that both vertices of each pair belong to the same path. For this problem, we show that, while it is NP-hard to compute if there exists a solution consisting of at most three paths, it is possible to decide in polynomial time whether a solution consisting of at most two paths exists. The third variant (MaxRPSP), given a DAG and a set of pairs of vertices, asks for a single path containing the maximum number of the given pairs of vertices. We show that MaxRPSP is W[1]-hard when parameterized by the number of covered pairs and we give a fixed-parameter algorithm when the parameter is the maximum overlapping degree.
Niko Beerenwinkel, Stefano Beretta 0001, Paola Bonizzoni, Riccardo Dondi, Yuri Pirola
Comput. J.1
2015 Identification of Constrained Cancer Driver Genes Based on Mutation Timing
abstract
Cancer drivers are genomic alterations that provide cells containing them with a selective advantage over their local competitors, whereas neutral passengers do not change the somatic fitness of cells. Cancer-driving mutations are usually discriminated from passenger mutations by their higher degree of recurrence in tumor samples. However, there is increasing evidence that many additional driver mutations may exist that occur at very low frequencies among tumors. This observation has prompted alternative methods for driver detection, including finding groups of mutually exclusive mutations and incorporating prior biological knowledge about gene function or network structure. Dependencies among drivers due to epistatic interactions can also result in low mutation frequencies, but this effect has been ignored in driver detection so far. Here, we present a new computational approach for identifying genomic alterations that occur at low frequencies because they depend on other events. Unlike passengers, these constrained mutations display punctuated patterns of occurrence in time. We test this driver-passenger discrimination approach based on mutation timing in extensive simulation studies, and we apply it to cross-sectional copy number alteration (CNA) data from ovarian cancer, CNA and single-nucleotide variant (SNV) data from breast tumors and SNV data from colorectal cancer. Among the top ranked predicted drivers, we find low-frequency genes that have already been shown to be involved in carcinogenesis, as well as many new candidate drivers. The mutation timing approach is orthogonal and complementary to existing driver prediction methods. It will help identifying from cancer genome data the alterations that drive tumor progression.
Thomas Sakoparnig, Patrick Fried, Niko Beerenwinkel
PLoS Comput. Biol.3
2015 NEMix: Single-cell Nested Effects Models for Probabilistic Pathway Stimulation
abstract
Nested effects models have been used successfully for learning subcellular networks from high-dimensional perturbation effects that result from RNA interference (RNAi) experiments. Here, we further develop the basic nested effects model using high-content single-cell imaging data from RNAi screens of cultured cells infected with human rhinovirus. RNAi screens with single-cell readouts are becoming increasingly common, and they often reveal high cell-to-cell variation. As a consequence of this cellular heterogeneity, knock-downs result in variable effects among cells and lead to weak average phenotypes on the cell population level. To address this confounding factor in network inference, we explicitly model the stimulation status of a signaling pathway in individual cells. We extend the framework of nested effects models to probabilistic combinatorial knock-downs and propose NEMix, a nested effects mixture model that accounts for unobserved pathway activation. We analyzed the identifiability of NEMix and developed a parameter inference scheme based on the Expectation Maximization algorithm. In an extensive simulation study, we show that NEMix improves learning of pathway structures over classical NEMs significantly in the presence of hidden pathway stimulation. We applied our model to single-cell imaging data from RNAi screens monitoring human rhinovirus infection, where limited infection efficiency of the assay results in uncertain pathway stimulation. Using a subset of genes with known interactions, we show that the inferred NEMix network has high accuracy and outperforms the classical nested effects model without hidden pathway activity. NEMix is implemented as part of the R/Bioconductor package 'nem' and available at www.cbg.ethz.ch/software/NEMix.
Juliane Siebourg-Polster, Daria Mudrak, Mario Emmenlauer, Pauli Rämö, Christoph Dehio, Urs F. Greber, Holger Fröhlich, Niko Beerenwinkel
PLoS Comput. Biol.8
2014 Covering Pairs in Directed Acyclic Graphs
Niko Beerenwinkel, Stefano Beretta 0001, Paola Bonizzoni, Riccardo Dondi, Yuri Pirola
LATA1
2014 Modeling Mutual Exclusivity of Cancer Mutations
Ewa Szczurek, Niko Beerenwinkel
RECOMB2
2014 Viral Quasispecies Assembly via Maximal Clique Enumeration
Armin Töpfer, Tobias Marschall, Rowena A. Bull, Fabio Luciani, Alexander Schönhuth, Niko Beerenwinkel
RECOMB6
2014 Exact likelihood computation in Boolean networks with probabilistic time delays, and its application in signal network reconstruction
abstract
MOTIVATION: For biological pathways, it is common to measure a gene expression time series after various knockdowns of genes that are putatively involved in the process of interest. These interventional time-resolved data are most suitable for the elucidation of dynamic causal relationships in signaling networks. Even with this kind of data it is still a major and largely unsolved challenge to infer the topology and interaction logic of the underlying regulatory network. RESULTS: In this work, we present a novel model-based approach involving Boolean networks to reconstruct small to medium-sized regulatory networks. In particular, we solve the problem of exact likelihood computation in Boolean networks with probabilistic exponential time delays. Simulations demonstrate the high accuracy of our approach. We apply our method to data of Ivanova et al. (2006), where RNA interference knockdown experiments were used to build a network of the key regulatory genes governing mouse stem cell maintenance and differentiation. In contrast to previous analyses of that data set, our method can identify feedback loops and provides new insights into the interplay of some master regulators in embryonic stem cell development. AVAILABILITY AND IMPLEMENTATION: The algorithm is implemented in the statistical language R. Code and documentation are available at Bioinformatics online. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary Materials are available at Bioinfomatics online.
Sebastian Dümcke, Johannes Bräuer, Benedict Anchang, Rainer Spang, Niko Beerenwinkel, Achim Tresch
Bioinform.5
2014 Challenges in RNA virus bioinformatics
abstract
MOTIVATION: Computer-assisted studies of structure, function and evolution of viruses remains a neglected area of research. The attention of bioinformaticians to this interesting and challenging field is far from commensurate with its medical and biotechnological importance. It is telling that out of >200 talks held at ISMB 2013, the largest international bioinformatics conference, only one presentation explicitly dealt with viruses. In contrast to many broad, established and well-organized bioinformatics communities (e.g. structural genomics, ontologies, next-generation sequencing, expression analysis), research groups focusing on viruses can probably be counted on the fingers of two hands. RESULTS: The purpose of this review is to increase awareness among bioinformatics researchers about the pressing needs and unsolved problems of computational virology. We focus primarily on RNA viruses that pose problems to many standard bioinformatics analyses owing to their compact genome organization, fast mutation rate and low evolutionary conservation. We provide an overview of tools and algorithms for handling viral sequencing data, detecting functionally important RNA structures, classifying viral proteins into families and investigating the origin and evolution of viruses.
Manja Marz, Niko Beerenwinkel, Christian Drosten, Markus Fricke, Dmitrij Frishman, Ivo L. Hofacker, Dieter Hoffmann, Martin Middendorf, Thomas Rattei, Peter F. Stadler, Armin Töpfer
Bioinform.2
2014 Estimating HIV-1 Fitness Characteristics from Cross-Sectional Genotype Data
abstract
Despite the success of highly active antiretroviral therapy (HAART) in the management of human immunodeficiency virus (HIV)-1 infection, virological failure due to drug resistance development remains a major challenge. Resistant mutants display reduced drug susceptibilities, but in the absence of drug, they generally have a lower fitness than the wild type, owing to a mutation-incurred cost. The interaction between these fitness costs and drug resistance dictates the appearance of mutants and influences viral suppression and therapeutic success. Assessing in vivo viral fitness is a challenging task and yet one that has significant clinical relevance. Here, we present a new computational modelling approach for estimating viral fitness that relies on common sparse cross-sectional clinical data by combining statistical approaches to learn drug-specific mutational pathways and resistance factors with viral dynamics models to represent the host-virus interaction and actions of drug mechanistically. We estimate in vivo fitness characteristics of mutant genotypes for two antiretroviral drugs, the reverse transcriptase inhibitor zidovudine (ZDV) and the protease inhibitor indinavir (IDV). Well-known features of HIV-1 fitness landscapes are recovered, both in the absence and presence of drugs. We quantify the complex interplay between fitness costs and resistance by computing selective advantages for different mutants. Our approach extends naturally to multiple drugs and we illustrate this by simulating a dual therapy with ZDV and IDV to assess therapy failure. The combined statistical and dynamical modelling approach may help in dissecting the effects of fitness costs and resistance with the ultimate aim of assisting the choice of salvage therapies after treatment failure.
Sathej Gopalakrishnan, Hesam Montazeri, Stephan Menz, Niko Beerenwinkel, Wilhelm Huisinga
PLoS Comput. Biol.4
2014 Modeling Mutual Exclusivity of Cancer Mutations
abstract
In large collections of tumor samples, it has been observed that sets of genes that are commonly involved in the same cancer pathways tend not to occur mutated together in the same patient.Such gene sets form mutually exclusive patterns of gene alterations in cancer genomic data.Computational approaches that detect mutually exclusive gene sets, rank and test candidate alteration patterns by rewarding the number of samples the pattern covers and by punishing its impurity, i.e., additional alterations that violate strict mutual exclusivity.However, the extant approaches do not account for possible observation errors.In practice, false negatives and especially false positives can severely bias evaluation and ranking of alteration patterns.To address these limitations, we develop a fully probabilistic, generative model of mutual exclusivity, explicitly taking coverage, impurity, as well as error rates into account, and devise efficient algorithms for parameter estimation and pattern ranking.Based on this model, we derive a statistical test of mutual exclusivity by comparing its likelihood to the null model that assumes independent gene alterations.Using extensive simulations, the new test is shown to be more powerful than a permutation test applied previously.When applied to detect mutual exclusivity patterns in glioblastoma and in pan-cancer data from twelve tumor types, we identify several significant patterns that are biologically relevant, most of which would not be detected by previous approaches.Our statistical modeling framework of mutual exclusivity provides increased flexibility and power to detect cancer pathways from genomic alteration data in the presence of noise.A summary of this paper appears in the proceedings of the RECOMB 2014 conference, April 2 5.
Ewa Szczurek, Niko Beerenwinkel
PLoS Comput. Biol.2
2014 Viral Quasispecies Assembly via Maximal Clique Enumeration
abstract
Virus populations can display high genetic diversity within individual hosts. The intra-host collection of viral haplotypes, called viral quasispecies, is an important determinant of virulence, pathogenesis, and treatment outcome. We present HaploClique, a computational approach to reconstruct the structure of a viral quasispecies from next-generation sequencing data as obtained from bulk sequencing of mixed virus samples. We develop a statistical model for paired-end reads accounting for mutations, insertions, and deletions. Using an iterative maximal clique enumeration approach, read pairs are assembled into haplotypes of increasing length, eventually enabling global haplotype assembly. The performance of our quasispecies assembly method is assessed on simulated data for varying population characteristics and sequencing technology parameters. Owing to its paired-end handling, HaploClique compares favorably to state-of-the-art haplotype inference methods. It can reconstruct error-free full-length haplotypes from low coverage samples and detect large insertions and deletions at low frequencies. We applied HaploClique to sequencing data derived from a clinical hepatitis C virus population of an infected patient and discovered a novel deletion of length 357±167 bp that was validated by two independent long-read sequencing experiments. HaploClique is available at https://github.com/armintoepfer/haploclique. A summary of this paper appears in the proceedings of the RECOMB 2014 conference, April 2-5.
Armin Töpfer, Tobias Marschall, Rowena A. Bull, Fabio Luciani, Alexander Schönhuth, Niko Beerenwinkel
PLoS Comput. Biol.6
2014 HIV Haplotype Inference Using a Propagating Dirichlet Process Mixture Model
abstract
This paper presents a new computational technique for the identification of HIV haplotypes. HIV tends to generate many potentially drug-resistant mutants within the HIV-infected patient and being able to identify these different mutants is important for efficient drug administration. With the view of identifying the mutants, we aim at analyzing short deep sequencing data called reads. From a statistical perspective, the analysis of such data can be regarded as a nonstandard clustering problem due to missing pairwise similarity measures between non-overlapping reads. To overcome this problem we propagate a Dirichlet Process Mixture Model by sequentially updating the prior information from successive local analyses. The model is verified using both simulated and real sequencing data.
Sandhya Prabhakaran, Mélanie Rey, Osvaldo Zagordi, Niko Beerenwinkel, Volker Roth 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2013 The Individualized Genetic Barrier Predicts Treatment Response in a Large Cohort of HIV-1 Infected Patients
abstract
The success of combination antiretroviral therapy is limited by the evolutionary escape dynamics of HIV-1. We used Isotonic Conjunctive Bayesian Networks (I-CBNs), a class of probabilistic graphical models, to describe this process. We employed partial order constraints among viral resistance mutations, which give rise to a limited set of mutational pathways, and we modeled phenotypic drug resistance as monotonically increasing along any escape pathway. Using this model, the individualized genetic barrier (IGB) to each drug is derived as the probability of the virus not acquiring additional mutations that confer resistance. Drug-specific IGBs were combined to obtain the IGB to an entire regimen, which quantifies the virus' genetic potential for developing drug resistance under combination therapy. The IGB was tested as a predictor of therapeutic outcome using between 2,185 and 2,631 treatment change episodes of subtype B infected patients from the Swiss HIV Cohort Study Database, a large observational cohort. Using logistic regression, significant univariate predictors included most of the 18 drugs and single-drug IGBs, the IGB to the entire regimen, the expert rules-based genotypic susceptibility score (GSS), several individual mutations, and the peak viral load before treatment change. In the multivariate analysis, the only genotype-derived variables that remained significantly associated with virological success were GSS and, with 10-fold stronger association, IGB to regimen. When predicting suppression of viral load below 400 cps/ml, IGB outperformed GSS and also improved GSS-containing predictors significantly, but the difference was not significant for suppression below 50 cps/ml. Thus, the IGB to regimen is a novel data-derived predictor of treatment outcome that has potential to improve the interpretation of genotypic drug resistance tests.
Niko Beerenwinkel, Hesam Montazeri, Heike Schuhmacher, Patrick Knupfer, Viktor von Wyl, Hansjakob Furrer, Manuel Battegay, Bernard Hirschel, Matthias Cavassini, Pietro Vernazza, Enos Bernasconi, Sabine Yerly, Jürg Böni, Thomas Klimkait, Cristina Cellerai, Huldrych F. Günthard
PLoS Comput. Biol.1
2012 Probabilistic Inference of Viral Quasispecies Subject to Recombination
Osvaldo Zagordi, Armin Töpfer, Sandhya Prabhakaran, Volker Roth 0001, Eran Halperin, Niko Beerenwinkel
RECOMB6
2012 Efficient sampling for Bayesian inference of conjunctive Bayesian networks
abstract
MOTIVATION: Cancer development is driven by the accumulation of advantageous mutations and subsequent clonal expansion of cells harbouring these mutations, but the order in which mutations occur remains poorly understood. Advances in genome sequencing and the soon-arriving flood of cancer genome data produced by large cancer sequencing consortia hold the promise to elucidate cancer progression. However, new computational methods are needed to analyse these large datasets. RESULTS: We present a Bayesian inference scheme for Conjunctive Bayesian Networks, a probabilistic graphical model in which mutations accumulate according to partial order constraints and cancer genotypes are observed subject to measurement noise. We develop an efficient MCMC sampling scheme specifically designed to overcome local optima induced by dependency structures. We demonstrate the performance advantage of our sampler over traditional approaches on simulated data and show the advantages of adopting a Bayesian perspective when reanalyzing cancer datasets and comparing our results to previous maximum-likelihood-based approaches. AVAILABILITY: An R package including the sampler and examples is available at http://www.cbg.ethz.ch/software/bayes-cbn. CONTACTS: [email protected].
Thomas Sakoparnig, Niko Beerenwinkel
Bioinform.2
2012 Stability of gene rankings from RNAi screens
abstract
MOTIVATION: Genome-wide RNA interference (RNAi) experiments are becoming a widely used approach for identifying intracellular molecular pathways of specific functions. However, detecting all relevant genes involved in a biological process is challenging, because typically only few samples per gene knock-down are available and readouts tend to be very noisy. We investigate the reliability of top scoring hit lists obtained from RNAi screens, compare the performance of different ranking methods, and propose a new ranking method to improve the reproducibility of gene selection. RESULTS: The performance of different ranking methods is assessed by the size of the stable sets they produce, i.e. the subsets of genes which are estimated to be re-selected with high probability in independent validation experiments. Using stability selection, we also define a new ranking method, called stability ranking, to improve the stability of any given base ranking method. Ranking methods based on mean, median, t-test and rank-sum test, and their stability-augmented counterparts are compared in simulation studies and on three microscopy image RNAi datasets. We find that the rank-sum test offers the most favorable trade-off between ranking stability and accuracy and that stability ranking improves the reproducibility of all and the accuracy of several ranking methods. AVAILABILITY: Stability ranking is freely available as the R/Bioconductor package staRank at http://www.cbg.ethz.ch/software/staRank.
Juliane Siebourg-Polster, Gunter Merdes, Benjamin Misselwitz, Wolf-Dietrich Hardt, Niko Beerenwinkel
Bioinform.5
2011 ShoRAH: estimating the genetic diversity of a mixed sample from next-generation sequencing data
abstract
BACKGROUND: With next-generation sequencing technologies, experiments that were considered prohibitive only a few years ago are now possible. However, while these technologies have the ability to produce enormous volumes of data, the sequence reads are prone to error. This poses fundamental hurdles when genetic diversity is investigated. RESULTS: We developed ShoRAH, a computational method for quantifying genetic diversity in a mixed sample and for identifying the individual clones in the population, while accounting for sequencing errors. The software was run on simulated data and on real data obtained in wet lab experiments to assess its reliability. CONCLUSIONS: ShoRAH is implemented in C++, Python, and Perl and has been tested under Linux and Mac OS X. Source code is available under the GNU General Public License at http://www.cbg.ethz.ch/software/shorah.
Osvaldo Zagordi, Arnab Bhattacharya 0003, Nicholas Eriksson, Niko Beerenwinkel
BMC Bioinform.4
2009 Deep Sequencing of a Genetically Heterogeneous Sample: Local Haplotype Reconstruction and Read Error Correction
Osvaldo Zagordi, Lukas Geyrhofer, Volker Roth 0001, Niko Beerenwinkel
RECOMB4
2009 Quantifying cancer progression with conjunctive Bayesian networks
abstract
MOTIVATION: Cancer is an evolutionary process characterized by accumulating mutations. However, the precise timing and the order of genetic alterations that drive tumor progression remain enigmatic. RESULTS: We present a specific probabilistic graphical model for the accumulation of mutations and their interdependencies. The Bayesian network models cancer progression by an explicit unobservable accumulation process in time that is separated from the observable but error-prone detection of mutations. Model parameters are estimated by an Expectation-Maximization algorithm and the underlying interaction graph is obtained by a simulated annealing procedure. Applying this method to cytogenetic data for different cancer types, we find multiple complex oncogenetic pathways deviating substantially from simplified models, such as linear pathways or trees. We further demonstrate how the inferred progression dynamics can be used to improve genetics-based survival predictions which could support diagnostics and prognosis. AVAILABILITY: The software package ct-cbn is available under a GPL license on the web site cbg.ethz.ch/software/ct-cbn CONTACT: [email protected].
Moritz Gerstung, Michael Baudis, Holger Moch, Niko Beerenwinkel
Bioinform.4
2008 Viral Population Estimation Using Pyrosequencing
abstract
The diversity of virus populations within single infected hosts presents a major difficulty for the natural immune response as well as for vaccine design and antiviral drug therapy. Recently developed pyrophosphate-based sequencing technologies (pyrosequencing) can be used for quantifying this diversity by ultra-deep sequencing of virus samples. We present computational methods for the analysis of such sequence data and apply these techniques to pyrosequencing data obtained from HIV populations within patients harboring drug-resistant virus strains. Our main result is the estimation of the population structure of the sample from the pyrosequencing reads. This inference is based on a statistical approach to error correction, followed by a combinatorial algorithm for constructing a minimal set of haplotypes that explain the data. Using this set of explaining haplotypes, we apply a statistical model to infer the frequencies of the haplotypes in the population via an expectation-maximization (EM) algorithm. We demonstrate that pyrosequencing reads allow for effective population reconstruction by extensive simulations and by comparison to 165 sequences obtained directly from clonal sequencing of four independent, diverse HIV populations. Thus, pyrosequencing can be used for cost-effective estimation of the structure of virus populations, promising new insights into viral evolutionary dynamics and disease control strategies.
Nicholas Eriksson, Lior Pachter, Yumi Mitsuya, Soo-Yon Rhee, Baback Gharizadeh, Mostafa Ronaghi, Robert W. Shafer, Niko Beerenwinkel
PLoS Comput. Biol.9
2007 Genetic Progression and the Waiting Time to Cancer
abstract
Cancer results from genetic alterations that disturb the normal cooperative behavior of cells. Recent high-throughput genomic studies of cancer cells have shown that the mutational landscape of cancer is complex and that individual cancers may evolve through mutations in as many as 20 different cancer-associated genes. We use data published by Sjöblom et al. (2006) to develop a new mathematical model for the somatic evolution of colorectal cancers. We employ the Wright-Fisher process for exploring the basic parameters of this evolutionary process and derive an analytical approximation for the expected waiting time to the cancer phenotype. Our results highlight the relative importance of selection over both the size of the cell population at risk and the mutation rate. The model predicts that the observed genetic diversity of cancer genomes can arise under a normal mutation rate if the average selective advantage per mutation is on the order of 1%. Increased mutation rates due to genetic instability would allow even smaller selective advantages during tumorigenesis. The complexity of cancer progression can be understood as the result of multiple sequential mutations, each of which has a relatively small but positive effect on net cell growth.
Niko Beerenwinkel, Tibor Antal, David Dingli, Arne Traulsen, Kenneth W. Kinzler, Victor E. Velculescu, Bert Vogelstein, Martin A. Nowak
PLoS Comput. Biol.1
2006 Mutagenetic tree Fisher kernel improves prediction of HIV drug resistance from viral genotype
abstract
Starting with the work of Jaakkola and Haussler, a variety of approaches have been proposed for coupling domain-specific generative models with statistical learning methods. The link is established by a kernel function which provides a similarity measure based inherently on the underlying model. In computational biology, the full promise of this framework has rarely ever been exploited, as most kernels are derived from very generic models, such as sequence profiles or hidden Markov models. Here, we introduce the MTreeMix kernel, which is based on a generative model tailored to the underlying biological mechanism. Specifically, the kernel quantifies the similarity of evolutionary escape from antiviral drug pressure between two viral sequence samples. We compare this novel kernel to a standard, evolution-agnostic amino acid encoding in the prediction of HIV drug resistance from genotype, using support vector regression. The results show significant improvements in predictive performance across 17 anti-HIV drugs. Thus, in our study, the generative-discriminative paradigm is key to bridging the gap between population genetic modeling and clinical decision making.
Tobias Sing, Niko Beerenwinkel
NIPS2
2005 Characterization of Novel HIV Drug Resistance Mutations Using Clustering, Multidimensional Scaling and SVM-Based Feature Ranking
Tobias Sing, Valentina Svicher, Niko Beerenwinkel, Francesca Ceccherini-Silberstein, Martin Däumer, Rolf Kaiser, Hauke Walter, Klaus Korn, Daniel Hoffmann, Mark Oette, Jürgen K. Rockstroh, Gerd Fätkenheuer, Carlo-Federico Perno, Thomas Lengauer
PKDD3
2005 Mtreemix: a software package for learning and using mixture models of mutagenetic trees
abstract
SUMMARY: Mixture models of mutagenetic trees constitute a class of probabilistic models for describing evolutionary processes that are characterized by the accumulation of permanent genetic changes. They have been applied to model the accumulation of chromosomal gains and losses in tumor development and the development of drug resistance-associated mutations in the HIV genome.Mtreemix is a software package for estimating mutagenetic trees mixture models from observed cross-sectional data and for using these models for predictions. We provide programs for model fitting, model selection, simulation, likelihood computation and waiting time estimation. AVAILABILITY: Mtreemix, including source code, documentation, sample data files and precompiled Solaris and Linux binaries, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de/
Niko Beerenwinkel, Jörg Rahnenführer, Rolf Kaiser, Daniel Hoffmann, Joachim Selbig, Thomas Lengauer
Bioinform.1
2005 Computational methods for the design of effective therapies against drug resistant HIV strains
abstract
The development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data.
Niko Beerenwinkel, Tobias Sing, Thomas Lengauer, Jörg Rahnenführer, Kirsten Roomp, Igor Savenkov, Roman Fischer, Daniel Hoffmann, Joachim Selbig, Klaus Korn, Hauke Walter, Thomas Berg, Patrick Braun, Gerd Fätkenheuer, Mark Oette, Jürgen K. Rockstroh, Bernd Kupfer, Rolf Kaiser, Martin Däumer
Bioinform.1
2005 Estimating cancer survival and clinical outcome based on genetic tumor progression scores
abstract
MOTIVATION: In cancer research, prediction of time to death or relapse is important for a meaningful tumor classification and selecting appropriate therapies. Survival prognosis is typically based on clinical and histological parameters. There is increasing interest in identifying genetic markers that better capture the status of a tumor in order to improve on existing predictions. The accumulation of genetic alterations during tumor progression can be used for the assessment of the genetic status of the tumor. For modeling dependences between the genetic events, evolutionary tree models have been applied. RESULTS: Mixture models of oncogenetic trees provide a probabilistic framework for the estimation of typical pathogenetic routes. From these models we derive a genetic progression score (GPS) that estimates the genetic status of a tumor. GPS is calculated for glioblastoma patients from loss of heterozygosity measurements and for prostate cancer patients from comparative genomic hybridization measurements. Cox proportional hazard models are then fitted to observed survival times of glioblastoma patients and to times until PSA relapse following radical prostatectomy of prostate cancer patients. It turns out that the genetically defined GPS is predictive even after adjustment for classical clinical markers and thus can be considered a medically relevant prognostic factor. AVAILABILITY: Mtreemix, a software package for estimating tree mixture models, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de. The raw cancer datasets and R code for the analysis with Cox models are available upon request from the corresponding author.
Jörg Rahnenführer, Niko Beerenwinkel, Wolfgang A. Schulz, Christian Hartmann 0006, Andreas von Deimling, Bernd Wullich, Thomas Lengauer
Bioinform.2
2005 ROCR: visualizing classifier performance in R
abstract
UNLABELLED: ROCR is a package for evaluating and visualizing the performance of scoring classifiers in the statistical language R. It features over 25 performance measures that can be freely combined to create two-dimensional performance curves. Standard methods for investigating trade-offs between specific performance measures are available within a uniform framework, including receiver operating characteristic (ROC) graphs, precision/recall plots, lift charts and cost curves. ROCR integrates tightly with R's powerful graphics capabilities, thus allowing for highly adjustable plots. Being equipped with only three commands and reasonable default values for optional parameters, ROCR combines flexibility with ease of usage. AVAILABILITY: http://rocr.bioinf.mpi-sb.mpg.de. ROCR can be used under the terms of the GNU General Public License. Running within R, it is platform-independent. CONTACT: [email protected].
Tobias Sing, Oliver Sander, Niko Beerenwinkel, Thomas Lengauer
Bioinform.3
2004 Learning multiple evolutionary pathways from cross-sectional data
abstract
We introduce a mixture model of trees to describe evolutionary processes that are characterized by the accumulation of permanent genetic changes. The basic building block of the model is a directed weighted tree that generates a probability distribution on the set of all patterns of genetic events. We present an EM-like algorithm for learning a mixture model of K trees and show how to determine K with a maximum likelihood approach. As a case study we consider the accumulation of mutations in the HIV-1 reverse transcriptase that are associated with drug resistance. The fitted model is statistically validated as a density estimator and the stability of the model topology is analyzed. We obtain a generative probabilistic model for the development of drug resistance in HIV that agrees with biological knowledge. Further applications and extensions of the model are discussed.
Niko Beerenwinkel, Jörg Rahnenführer, Martin Däumer, Daniel Hoffmann, Rolf Kaiser, Joachim Selbig, Thomas Lengauer
RECOMB1