EDBT 2026 Demo / reviewers in the wild / expert
Jack Kuipers
dblp:138/6428
· DBLP profile ↗
22ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-5357-2705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Single-cell copy number calling and event history reconstructionabstractMOTIVATION: Copy number alterations are driving forces of tumour development and the emergence of intra-tumour heterogeneity. A comprehensive picture of these genomic aberrations is therefore essential for the development of personalised and precise cancer diagnostics and therapies. Single-cell sequencing offers the highest resolution for copy number profiling down to the level of individual cells. Recent high-throughput protocols allow for the processing of hundreds of cells through shallow whole-genome DNA sequencing. The resulting low read-depth data poses substantial statistical and computational challenges to the identification of copy number alterations. RESULTS: We developed SCICoNE, a statistical model and MCMC algorithm tailored to single-cell copy number profiling from shallow whole-genome DNA sequencing data. SCICoNE reconstructs the history of copy number events in the tumour and uses these evolutionary relationships to identify the copy number profiles of the individual cells. We show the accuracy of this approach in evaluations on simulated data and demonstrate its practicability in applications to two breast cancer samples from different sequencing protocols. AVAILABILITY AND IMPLEMENTATION: SCICoNE is available at https://github.com/cbg-ethz/SCICoNE. Jack Kuipers, Mustafa Anil Tuncel, Pedro F. Ferreira, Katharina Jahn 0001, Niko Beerenwinkel |
Bioinform. | 1 |
| 2025 | Bayesian inference of fitness landscapes via tree-structured branching processesabstractMOTIVATION: The complex dynamics of cancer evolution, driven by mutation and selection, underlies the molecular heterogeneity observed in tumors. The evolutionary histories of tumors of different patients can be encoded as mutation trees and reconstructed in high resolution from single-cell sequencing data, offering crucial insights for studying fitness effects of and epistasis among mutations. Existing models, however, either fail to separate mutation and selection or neglect the evolutionary histories encoded by the tumor phylogenetic trees. RESULTS: We introduce FiTree, a tree-structured multi-type branching process model with epistatic fitness parameterization and a Bayesian inference scheme to learn fitness landscapes from single-cell tumor mutation trees. Through simulations, we demonstrate that FiTree outperforms state-of-the-art methods in inferring the fitness landscape underlying tumor evolution. Applying FiTree to a single-cell acute myeloid leukemia dataset, we identify epistatic fitness effects consistent with known biological findings and quantify uncertainty in predicting future mutational events. The new model unifies probabilistic graphical models of cancer progression with population genetics, offering a principled framework for understanding tumor evolution and informing therapeutic strategies. AVAILABILITY AND IMPLEMENTATION: The Python package FiTree and the analysis workflows are available at https://github.com/cbg-ethz/FiTree. Xiang Ge Luo, Jack Kuipers, Kevin Rupp, Koichi Takahashi, Niko Beerenwinkel |
Bioinform. | 2 |
| 2023 | A Bayesian Take on Gaussian Process NetworksabstractGaussian Process Networks (GPNs) are a class of directed graphical models which employ Gaussian processes as priors for the conditional expectation of each variable given its parents in the network. The model allows the description of continuous joint distributions in a compact but flexible manner with minimal parametric assumptions on the dependencies between variables. Bayesian structure learning of GPNs requires computing the posterior over graphs of the network and is computationally infeasible even in low dimensions. This work implements Monte Carlo and Markov Chain Monte Carlo methods to sample from the posterior distribution of network structures. As such, the approach follows the Bayesian paradigm, comparing models via their marginal likelihood and computing the posterior probability of the GPN features. Simulation studies show that our method outperforms state-of-the-art algorithms in recovering the graphical structure of the network and provides an accurate approximation of its posterior distribution. Enrico Giudice, Jack Kuipers, Giusi Moffa |
NeurIPS | 2 |
| 2023 | The dual PC algorithm and the role of Gaussianity for structure learning of Bayesian networksabstractLearning the graphical structure of Bayesian networks is key to describing data-generating mechanisms in many complex applications but poses considerable computational challenges. Observational data can only identify the equivalence class of the directed acyclic graph underlying a Bayesian network model, and a variety of methods exist to tackle the problem. Under certain assumptions, the popular PC algorithm can consistently recover the correct equivalence class by reverse-engineering the conditional independence (CI) relationships holding in the variable distribution. The dual PC algorithm is a novel scheme to carry out the CI tests within the PC algorithm by leveraging the inverse relationship between covariance and precision matrices. By exploiting block matrix inversions we can also perform tests on partial correlations of complementary (or dual) conditioning sets. The multiple CI tests of the dual PC algorithm proceed by first considering marginal and full-order CI relationships and progressively moving to central-order ones. Simulation studies show that the dual PC algorithm outperforms the classic PC algorithm both in terms of run time and in recovering the underlying network structure, even in the presence of deviations from Gaussianity. Additionally, we show that the dual PC algorithm applies for Gaussian copula models, and demonstrate its performance in that setting. Enrico Giudice, Jack Kuipers, Giusi Moffa |
Int. J. Approx. Reason. | 2 |
| 2022 | Mapping Single-Cell Transcriptomes to Copy Number Evolutionary Trees
Pedro F. Ferreira, Jack Kuipers, Niko Beerenwinkel |
RECOMB | 2 |
| 2022 | Joint Inference of Repeated Evolutionary Trajectories and Patterns of Clonal Exclusivity or Co-occurrence from Tumor Mutation Trees
Xiang Ge Luo, Jack Kuipers, Niko Beerenwinkel |
RECOMB | 2 |
| 2022 | Discovering gene regulatory networks of multiple phenotypic groups using dynamic Bayesian networksabstractDynamic Bayesian networks (DBNs) can be used for the discovery of gene regulatory networks (GRNs) from time series gene expression data. Here, we suggest a strategy for learning DBNs from gene expression data by employing a Bayesian approach that is scalable to large networks and is targeted at learning models with high predictive accuracy. Our framework can be used to learn DBNs for multiple groups of samples and highlight differences and similarities in their GRNs. We learn these DBN models based on different structural and parametric assumptions and select the optimal model based on the cross-validated predictive accuracy. We show in simulation studies that our approach is better equipped to prevent overfitting than techniques used in previous studies. We applied the proposed DBN-based approach to two time series transcriptomic datasets from the Gene Expression Omnibus database, each comprising data from distinct phenotypic groups of the same tissue type. In the first case, we used DBNs to characterize responders and non-responders to anti-cancer therapy. In the second case, we compared normal to tumor cells of colorectal tissue. The classification accuracy reached by the DBN-based classifier for both datasets was higher than reported previously. For the colorectal cancer dataset, our analysis suggested that GRNs for cancer and normal tissues have a lot of differences, which are most pronounced in the neighborhoods of oncogenes and known cancer tissue markers. The identified differences in gene networks of cancer and normal cells may be used for the discovery of targeted therapies. Polina Suter, Jack Kuipers, Niko Beerenwinkel |
Briefings Bioinform. | 2 |
| 2022 | Single-cell mutation calling and phylogenetic tree reconstruction with loss and recurrenceabstractMOTIVATION: Tumours evolve as heterogeneous populations of cells, which may be distinguished by different genomic aberrations. The resulting intra-tumour heterogeneity plays an important role in cancer patient relapse and treatment failure, so that obtaining a clear understanding of each patient's tumour composition and evolutionary history is key for personalized therapies. Single-cell sequencing (SCS) now provides the possibility to resolve tumour heterogeneity at the highest resolution of individual tumour cells, but brings with it challenges related to the particular noise profiles of the sequencing protocols as well as the complexity of the underlying evolutionary process. RESULTS: By modelling the noise processes and allowing mutations to be lost or to reoccur during tumour evolution, we present a method to jointly call mutations in each cell, reconstruct the phylogenetic relationship between cells, and determine the locations of mutational losses and recurrences. Our Bayesian approach allows us to accurately call mutations as well as to quantify our certainty in such predictions. We show the advantages of allowing mutational loss or recurrence with simulated data and present its application to tumour SCS data. AVAILABILITY AND IMPLEMENTATION: SCIΦN is available at https://github.com/cbg-ethz/SCIPhIN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jack Kuipers, Jochen Singer, Niko Beerenwinkel |
Bioinform. | 1 |
| 2022 | scAmpi - A versatile pipeline for single-cell RNA-seq analysis from basics to clinicsabstractSingle-cell RNA sequencing (scRNA-seq) has emerged as a powerful technique to decipher tissue composition at the single-cell level and to inform on disease mechanisms, tumor heterogeneity, and the state of the immune microenvironment. Although multiple methods for the computational analysis of scRNA-seq data exist, their application in a clinical setting demands standardized and reproducible workflows, targeted to extract, condense, and display the clinically relevant information. To this end, we designed scAmpi (Single Cell Analysis mRNA pipeline), a workflow that facilitates scRNA-seq analysis from raw read processing to informing on sample composition, clinically relevant gene and pathway alterations, and in silico identification of personalized candidate drug treatments. We demonstrate the value of this workflow for clinical decision making in a molecular tumor board as part of a clinical study. Anne Bertolini, Michael Prummer, Mustafa Anil Tuncel, Ulrike Menzel, María Lourdes Rosano-González, Jack Kuipers, Daniel J. Stekhoven, Niko Beerenwinkel, Franziska Singer |
PLoS Comput. Biol. | 6 |
| 2022 | Multi-omics subtyping of hepatocellular carcinoma patients using a Bayesian network mixture modelabstractComprehensive molecular characterization of cancer subtypes is essential for predicting clinical outcomes and searching for personalized treatments. We present bnClustOmics, a statistical model and computational tool for multi-omics unsupervised clustering, which serves a dual purpose: Clustering patient samples based on a Bayesian network mixture model and learning the networks of omics variables representing these clusters. The discovered networks encode interactions among all omics variables and provide a molecular characterization of each patient subgroup. We conducted simulation studies that demonstrated the advantages of our approach compared to other clustering methods in the case where the generative model is a mixture of Bayesian networks. We applied bnClustOmics to a hepatocellular carcinoma (HCC) dataset comprising genome (mutation and copy number), transcriptome, proteome, and phosphoproteome data. We identified three main HCC subtypes together with molecular characteristics, some of which are associated with survival even when adjusting for the clinical stage. Cluster-specific networks shed light on the links between genotypes and molecular phenotypes of samples within their respective clusters and suggest targets for personalized treatments. Polina Suter, Eva Dazert, Jack Kuipers, Charlotte K. Y. Ng, Tuyana Boldanova, Michael N. Hall, Markus H. Heim, Niko Beerenwinkel |
PLoS Comput. Biol. | 3 |
| 2021 | Learning Bayesian Networks from Ordinal DataabstractBayesian networks are a powerful framework for studying the dependency structure of variables in a complex system. The problem of learning Bayesian networks is tightly associated with the given data type. Ordinal data, such as stages of cancer, rating scale survey questions, and letter grades for exams, are ubiquitous in applied research. However, existing solutions are mainly for continuous and nominal data. In this work, we propose an iterative score-and-search method - called the Ordinal Structural EM (OSEM) algorithm - for learning Bayesian networks from ordinal data. Unlike traditional approaches designed for nominal data, we explicitly respect the ordering amongst the categories. More precisely, we assume that the ordinal variables originate from marginally discretizing a set of Gaussian variables, whose structural dependence in the latent space follows a directed acyclic graph. Then, we adopt the Structural EM algorithm and derive closed-form scoring functions for efficient graph searching. Through simulation studies, we illustrate the superior performance of the OSEM algorithm compared to the alternatives and analyze various factors that may influence the learning accuracy. Finally, we demonstrate the practicality of our method with a real-world application on psychological survey data from 408 patients with co-morbid symptoms of obsessive-compulsive disorder and depression. Xiang Ge Luo, Giusi Moffa, Jack Kuipers |
J. Mach. Learn. Res. | 3 |
| 2021 | Statistical tests for intra-tumour clonal co-occurrence and exclusivityabstractTumour progression is an evolutionary process in which different clones evolve over time, leading to intra-tumour heterogeneity. Interactions between clones can affect tumour evolution and hence disease progression and treatment outcome. Intra-tumoural pairs of mutations that are overrepresented in a co-occurring or clonally exclusive fashion over a cohort of patient samples may be suggestive of a synergistic effect between the different clones carrying these mutations. We therefore developed a novel statistical testing framework, called GeneAccord, to identify such gene pairs that are altered in distinct subclones of the same tumour. We analysed our framework for calibration and power. By comparing its performance to baseline methods, we demonstrate that to control type I errors, it is essential to account for the evolutionary dependencies among clones. In applying GeneAccord to the single-cell sequencing of a cohort of 123 acute myeloid leukaemia patients, we find 1 clonally co-occurring and 8 clonally exclusive gene pairs. The clonally exclusive pairs mostly involve genes of the key signalling pathways. Jack Kuipers, Ariane L. Moore, Katharina Jahn 0001, Peter Schraml, Kiyomi Morita, P. Andrew Futreal, Koichi Takahashi, Christian Beisel, Holger Moch, Niko Beerenwinkel |
PLoS Comput. Biol. | 1 |
| 2021 | Comparing mutational pathways to lopinavir resistance in HIV-1 subtypes B versus CabstractAlthough combination antiretroviral therapies seem to be effective at controlling HIV-1 infections regardless of the viral subtype, there is increasing evidence for subtype-specific drug resistance mutations. The order and rates at which resistance mutations accumulate in different subtypes also remain poorly understood. Most of this knowledge is derived from studies of subtype B genotypes, despite not being the most abundant subtype worldwide. Here, we present a methodology for the comparison of mutational networks in different HIV-1 subtypes, based on Hidden Conjunctive Bayesian Networks (H-CBN), a probabilistic model for inferring mutational networks from cross-sectional genotype data. We introduce a Monte Carlo sampling scheme for learning H-CBN models for a larger number of resistance mutations and develop a statistical test to assess differences in the inferred mutational networks between two groups. We apply this method to infer the temporal progression of mutations conferring resistance to the protease inhibitor lopinavir in a large cross-sectional cohort of HIV-1 subtype C genotypes from South Africa, as well as to a data set of subtype B genotypes obtained from the Stanford HIV Drug Resistance Database and the Swiss HIV Cohort Study. We find strong support for different initial mutational events in the protease, namely at residue 46 in subtype B and at residue 82 in subtype C. The inferred mutational networks for subtype B versus C are significantly different sharing only five constraints on the order of accumulating mutations with mutation at residue 54 as the parental event. The results also suggest that mutations can accumulate along various alternative paths within subtypes, as opposed to a unique total temporal ordering. Beyond HIV drug resistance, the statistical methodology is applicable more generally for the comparison of inferred mutational networks between any two groups. Susana Posada-Céspedes, Gert van Zyl, Hesam Montazeri, Jack Kuipers, Soo-Yon Rhee, Roger D. Kouyos, Huldrych F. Günthard, Niko Beerenwinkel |
PLoS Comput. Biol. | 4 |
| 2020 | Predicting colorectal cancer risk from adenoma detection via a two-type branching process modelabstractDespite advances in the modeling and understanding of colorectal cancer development, the dynamics of the progression from benign adenomatous polyp to colorectal carcinoma are still not fully resolved.To take advantage of adenoma size and prevalence data in the National Endoscopic Database of the Clinical Outcomes Research Initiative (CORI) as well as colorectal cancer incidence and size data from the Surveillance Epidemiology and End Results (SEER) database, we construct a two-type branching process model with compartments representing adenoma and carcinoma cells.To perform parameter inference we present a new large-size approximation to the size distribution of the cancer compartment and validate our approach on simulated data.By fitting the model to the CORI and SEER data, we learn biologically relevant parameters, including the transition rate from adenoma to cancer.The inferred parameters allow us to predict the individualized risk of the presence of cancer cells for each screened patient.We provide a web application which allows the user to calculate these individual probabilities at https://ccrc-eth.shinyapps.io/CCRC/.For example, we find a 1 in 100 chance of cancer given the presence of an adenoma between 10 and 20mm size in an average risk patient at age 50.We show that our two-type branching process model recapitulates the early growth dynamics of colon adenomas and cancers and can recover epidemiological trends such as adenoma prevalence and cancer incidence while remaining mathematically and computationally tractable. Author summaryColorectal cancer is a major public health burden.The development of colorectal cancer starts with the mutational initiation of non-cancerous growths in the form of benign adenomatous polyps.These adenomas grow over time with the potential to develop into carcinomas.Many mathematical and simulation-based models have been used to gain insight into this process.We aimed to understand rates of adenoma growth and transition Brian Lang, Jack Kuipers, Benjamin Misselwitz, Niko Beerenwinkel |
PLoS Comput. Biol. | 2 |
| 2018 | Integrative Inference of Subclonal Tumour Evolution from Single-Cell and Bulk Sequencing Data
Salem Malikic, Katharina Jahn 0001, Jack Kuipers, Süleyman Cenk Sahinalp, Niko Beerenwinkel |
RECOMB | 3 |
| 2018 | SCIΦ: Single-Cell Mutation Identification via Phylogenetic Inference
Jochen Singer, Jack Kuipers, Katharina Jahn 0001, Niko Beerenwinkel |
RECOMB | 2 |
| 2018 | Improved pathway reconstruction from RNA interference screens by exploiting off-target effectsabstractMotivation: Pathway reconstruction has proven to be an indispensable tool for analyzing the molecular mechanisms of signal transduction underlying cell function. Nested effects models (NEMs) are a class of probabilistic graphical models designed to reconstruct signalling pathways from high-dimensional observations resulting from perturbation experiments, such as RNA interference (RNAi). NEMs assume that the short interfering RNAs (siRNAs) designed to knockdown specific genes are always on-target. However, it has been shown that most siRNAs exhibit strong off-target effects, which further confound the data, resulting in unreliable reconstruction of networks by NEMs. Results: Here, we present an extension of NEMs called probabilistic combinatorial nested effects models (pc-NEMs), which capitalize on the ancillary siRNA off-target effects for network reconstruction from combinatorial gene knockdown data. Our model employs an adaptive simulated annealing search algorithm for simultaneous inference of network structure and error rates inherent to the data. Evaluation of pc-NEMs on simulated data with varying number of phenotypic effects and noise levels as well as real data demonstrates improved reconstruction compared to classical NEMs. Application to Bartonella henselae infection RNAi screening data yielded an eight node network largely in agreement with previous works, and revealed novel binary interactions of direct impact between established components. Availability and implementation: The software used for the analysis is freely available as an R package at https://github.com/cbg-ethz/pcNEM.git. Supplementary information: Supplementary data are available at Bioinformatics online. Sumana Srivatsa, Jack Kuipers, Fabian Schmich, Simone Eicher, Mario Emmenlauer, Christoph Dehio, Niko Beerenwinkel |
Bioinform. | 2 |
| 2017 | Detailed simulation of cancer exome sequencing data reveals differences and common limitations of variant callersabstractBACKGROUND: Next-generation sequencing of matched tumor and normal biopsy pairs has become a technology of paramount importance for precision cancer treatment. Sequencing costs have dropped tremendously, allowing the sequencing of the whole exome of tumors for just a fraction of the total treatment costs. However, clinicians and scientists cannot take full advantage of the generated data because the accuracy of analysis pipelines is limited. This particularly concerns the reliable identification of subclonal mutations in a cancer tissue sample with very low frequencies, which may be clinically relevant. RESULTS: Using simulations based on kidney tumor data, we compared the performance of nine state-of-the-art variant callers, namely deepSNV, GATK HaplotypeCaller, GATK UnifiedGenotyper, JointSNVMix2, MuTect, SAMtools, SiNVICT, SomaticSniper, and VarScan2. The comparison was done as a function of variant allele frequencies and coverage. Our analysis revealed that deepSNV and JointSNVMix2 perform very well, especially in the low-frequency range. We attributed false positive and false negative calls of the nine tools to specific error sources and assigned them to processing steps of the pipeline. All of these errors can be expected to occur in real data sets. We found that modifying certain steps of the pipeline or parameters of the tools can lead to substantial improvements in performance. Furthermore, a novel integration strategy that combines the ranks of the variants yielded the best performance. More precisely, the rank-combination of deepSNV, JointSNVMix2, MuTect, SiNVICT and VarScan2 reached a sensitivity of 78% when fixing the precision at 90%, and outperformed all individual tools, where the maximum sensitivity was 71% with the same precision. CONCLUSIONS: The choice of well-performing tools for alignment and variant calling is crucial for the correct interpretation of exome sequencing data obtained from mixed samples, and common pipelines are suboptimal. We were able to relate observed substantial differences in performance to the underlying statistical models of the tools, and to pinpoint the error sources of false positive and false negative calls. These findings might inspire new software developments that improve exome sequencing pipelines and further the field of precision cancer treatment. Ariane L. Hofmann, Jonas Behr, Jochen Singer, Jack Kuipers, Christian Beisel, Peter Schraml, Holger Moch, Niko Beerenwinkel |
BMC Bioinform. | 4 |
| 2017 | Multistate nested canalizing functions and their networks
Claus Kadelka, Jack Kuipers, John O. Adeyeye, Reinhard C. Laubenbacher |
Theor. Comput. Sci. | 3 |
| 2016 | pathTiMEx: Joint Inference of Mutually Exclusive Cancer Pathways and Their Dependencies in Tumor Progression
Simona Cristea, Jack Kuipers, Niko Beerenwinkel |
RECOMB | 2 |
| 2016 | Tree Inference for Single-Cell Data
Katharina Jahn 0001, Jack Kuipers, Niko Beerenwinkel |
RECOMB | 2 |
| 2016 | Large-scale inference of conjunctive Bayesian networksabstractUNLABELLED: The continuous time conjunctive Bayesian network (CT-CBN) is a graphical model for analyzing the waiting time process of the accumulation of genetic changes (mutations). CT-CBN models have been successfully used in several biological applications such as HIV drug resistance development and genetic progression of cancer. However, current approaches for parameter estimation and network structure learning of CBNs can only deal with a small number of mutations (<20). Here, we address this limitation by presenting an efficient and accurate approximate inference algorithm using a Monte Carlo expectation-maximization algorithm based on importance sampling. The new method can now be used for a large number of mutations, up to one thousand, an increase by two orders of magnitude. In simulation studies, we present the accuracy as well as the running time efficiency of the new inference method and compare it with a MLE method, expectation-maximization, and discrete time CBN model, i.e. a first-order approximation of the CT-CBN model. We also study the application of the new model on HIV drug resistance datasets for the combination therapy with zidovudine plus lamivudine (AZT + 3TC) as well as under no treatment, both extracted from the Swiss HIV Cohort Study database. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented as an R package available at https://github.com/cbg-ethz/MC-CBN CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hesam Montazeri, Jack Kuipers, Roger D. Kouyos, Jürg Böni, Sabine Yerly, Thomas Klimkait, Vincent Aubert, Huldrych F. Günthard, Niko Beerenwinkel |
Bioinform. | 2 |