VLDB 2026 Research / reviewers in the wild / expert
Johannes Köster
dblp:118/6626
· DBLP profile ↗
16ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0001-9818-9320ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alignoth : portable and interactive visualization of read alignmentsabstractSUMMARY: We present Alignoth, a lightweight command line application that generates self-contained portable HTML reports of DNA sequencing read alignment pileups and additionally supports export to static formats such as PNG, SVG, and PDF as well as a JSON based embeddable representation. The HTML reports feature read name search and mapping-quality-based read highlighting, and require only minimal storage, making them practical to share, inspect, or integrate into broader reporting systems. They can be created in headless (i.e. terminal only) environments while being interactively inspected afterwards. AVAILABILITY AND IMPLEMENTATION: Alignoth is freely available under the MIT license at https://github.com/alignoth/alignoth (doi: https://doi.org/10.5281/zenodo.15837719). It is implemented in Rust and can be installed via Cargo or Conda. Felix Wiegand, Felix Mölder, Johannes Köster |
Bioinform. | 3 |
| 2026 | Virus variant quantification with OrthanqabstractBACKGROUND: Existing tools for virus variant identification can pinpoint the most abundant virus variant in a sequencing sample. However, patients can be infected by more than one variant of the same virus species or strain, for example by multiple variants of SARS-CoV-2. This leads to the more complicated problem of virus variant quantification from samples containing virus mixtures. RESULTS: We report on improvements of Orthanq, our generic tool for haplotype quantification, and show how it can be applied to perform uncertainty aware quantification of virus variants. We evaluate this ability on simulated and real SARS-CoV-2 and HIV-1 virus mixture datasets and show that Orthanq outperforms other state of the art approaches. CONCLUSIONS: Orthanq performs identification and uncertainty-aware quantification of known virus variants of any virus species, in particular also in samples with mixed infections. With extensive built-in visualizations and reporting of alternative solutions with posterior densities, users can easily evaluate the uncertainty of the results. Hamdiye Uzuner, Felix Wiegand, Sven Schrinner, David Laehnemann, Dirk Schadendorf, Johannes Köster |
BMC Bioinform. | 6 |
| 2026 | A terminology for scientific workflow systems
Frédéric Suter, Tainã Coleman, Ilkay Altintas, Rosa M. Badia, Bartosz Balis, Kyle Chard, Iacopo Colonnelli, Ewa Deelman, Paolo Di Tommaso, Thomas Fahringer, Carole A. Goble, Shantenu Jha, Daniel S. Katz, Johannes Köster, Ulf Leser, Kshitij Mehta, Hilary Oliver, Jayson Luc Peterson, Giovanni Pizzi, Loïc Pottier, Raül Sirvent, Eric Suchyta, Douglas Thain, Sean R. Wilkinson, Justin M. Wozniak, Rafael Ferreira da Silva |
Future Gener. Comput. Syst. | 14 |
| 2024 | Orthanq: transparent and uncertainty-aware haplotype quantification with application in HLA-typingabstractBACKGROUND: Identification of human leukocyte antigen (HLA) types from DNA-sequenced human samples is important in organ transplantation and cancer immunotherapy and remains a challenging task considering sequence homology and extreme polymorphism of HLA genes. RESULTS: We present Orthanq, a novel statistical model and corresponding application for transparent and uncertainty-aware quantification of haplotypes. We utilize our approach to perform HLA typing while, for the first time, reporting uncertainty of predictions and transparently observing mutations beyond reported HLA types. Using 99 gold standard samples from 1000 Genomes, Illumina Platinum Genomes and Genome In a Bottle projects, we show that Orthanq can provide overall superior accuracy and shorter runtimes than state-of-the-art HLA typers. CONCLUSIONS: Orthanq is the first approach that allows to directly utilize existing pangenome alignments and type all HLA loci. Moreover, it can be generalized for usages beyond HLA typing, e.g. for virus lineage quantification. Orthanq is available under https://orthanq.github.io . Hamdiye Uzuner, Annette Paschen, Dirk Schadendorf, Johannes Köster |
BMC Bioinform. | 4 |
| 2023 | Insane in the vembrane: filtering and transforming VCF/BCF filesabstractSUMMARY: We present vembrane as a command line variant call format (VCF)/binary call format (BCF) filtering tool that consolidates and extends the filtering functionality of previous software to meet any imaginable filtering use case. Vembrane exposes the VCF/BCF file type specification and its inofficial extensions by the annotation tools VEP and SnpEff as Python data structures. vembrane filter enables filtration by Python expressions, requiring only basic knowledge of the Python programming language. vembrane table allows users to generate tables from subsets of annotations or functions thereof. Finally, it is fast, by using pysam and relying on lazy evaluation. AVAILABILITY AND IMPLEMENTATION: Source code and installation instructions are available at github.com/vembrane/vembrane (doi: 10.5281/zenodo.7003981). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Till Hartmann, Christopher Schröder 0002, Elias Kuthe, David Laehnemann, Johannes Köster |
Bioinform. | 5 |
| 2021 | VISPR-online: a web-based interactive tool to visualize CRISPR screening experimentsabstractBACKGROUND: VISPR is an interactive visualization and analysis framework for CRISPR screening experiments. However, it only supports the output of MAGeCK, and requires installation and manual configuration. Furthermore, VISPR is designed to run on a single computer, and data sharing between collaborators is challenging. RESULTS: To make the tool easily accessible to the community, we present VISPR-online, a web-based general application allowing users to visualize, explore, and share CRISPR screening data online with a few simple steps. VISPR-online provides an exploration of screening results and visualization of read count changes. Apart from MAGeCK, VISPR-online supports two more popular CRISPR screening analysis tools: BAGEL and JACKS. It provides an interactive environment for exploring gene essentiality, viewing guide RNA (gRNA) locations, and allowing users to resume and share screening results. CONCLUSIONS: VISPR-online allows users to visualize, explore and share CRISPR screening data online. It is freely available at http://vispr-online.weililab.org , while the source code is available at https://github.com/lemoncyb/VISPR-online . Yingbo Cui 0001, Johannes Köster, Xiangke Liao, Shaoliang Peng, Tao Tang 0001, Chun Huang 0006, Canqun Yang |
BMC Bioinform. | 3 |
| 2019 | Protein Complex Similarity Based on Weisfeiler-Lehman Labeling
Bianca K. Stöcker, Till Schäfer, Petra Mutzel, Johannes Köster, Nils M. Kriege, Sven Rahmann |
SISAP | 4 |
| 2019 | Full-length de novo viral quasispecies assembly through variation graph constructionabstractMOTIVATION: Viruses populate their hosts as a viral quasispecies: a collection of genetically related mutant strains. Viral quasispecies assembly is the reconstruction of strain-specific haplotypes from read data, and predicting their relative abundances within the mix of strains is an important step for various treatment-related reasons. Reference genome independent ('de novo') approaches have yielded benefits over reference-guided approaches, because reference-induced biases can become overwhelming when dealing with divergent strains. While being very accurate, extant de novo methods only yield rather short contigs. The remaining challenge is to reconstruct full-length haplotypes together with their abundances from such contigs. RESULTS: We present Virus-VG as a de novo approach to viral haplotype reconstruction from preassembled contigs. Our method constructs a variation graph from the short input contigs without making use of a reference genome. Then, to obtain paths through the variation graph that reflect the original haplotypes, we solve a minimization problem that yields a selection of maximal-length paths that is, optimal in terms of being compatible with the read coverages computed for the nodes of the variation graph. We output the resulting selection of maximal length paths as the haplotypes, together with their abundances. Benchmarking experiments on challenging simulated and real datasets show significant improvements in assembly contiguity compared to the input contigs, while preserving low error rates compared to the state-of-the-art viral quasispecies assemblers. AVAILABILITY AND IMPLEMENTATION: Virus-VG is freely available at https://bitbucket.org/jbaaijens/virus-vg. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jasmijn A. Baaijens, Bastiaan Van der Roest, Johannes Köster, Leen Stougie, Alexander Schönhuth |
Bioinform. | 3 |
| 2019 | A Bayesian model for single cell transcript expression analysis on MERFISH dataabstractMOTIVATION: Multiplexed error-robust fluorescence in-situ hybridization (MERFISH) is a recent technology to obtain spatially resolved gene or transcript expression profiles in single cells for hundreds to thousands of genes in parallel. So far, no statistical framework to analyze MERFISH data is available. RESULTS: We present a Bayesian model for single cell transcript expression analysis on MERFISH data. We show that the model successfully captures uncertainty in MERFISH data and eliminates systematic biases that can occur in raw RNA molecule counts obtained with MERFISH. Our model accurately estimates transcript expression and additionally provides the full probability distribution and credible intervals for each transcript. We further show how this enables MERFISH to scale towards the whole genome while being able to control the uncertainty in obtained results. AVAILABILITY AND IMPLEMENTATION: The presented model is implemented on top of Rust-Bio (Köster, 2016) and available open-source as MERFISHtools (https://merfishtools.github.io). It can be easily installed via Bioconda (Grüning et al., 2018). The entire analysis performed in this paper is provided as a fully reproducible Snakemake (Köster and Rahmann, 2012) workflow via Zenodo (https://doi.org/10.5281/zenodo.752340). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Johannes Köster, Myles Brown, Xiaole Shirley Liu |
Bioinform. | 1 |
| 2019 | A Bayesian model for single cell transcript expression analysis on MERFISH dataabstractBioinformatics, (2018) doi.org/10.1093/bioinformatics/bty718 The authors of the above paper wish to inform the reader that an affiliation for Johannes Köster was not noted in the original publication and also that the ‘Availability and Implementation’ section of the abstract should have appeared as below: ‘The presented model is implemented on top of Rust-Bio (Köster, 2016) and available open-source as MERFISHtools (https://merfishtools.github.io). It can be easily installed via Bioconda (Grüning et al., 2018). The entire analysis performed in this paper is provided as a fully reproducible Snakemake (Köster and Rahmann, 2012) workflow via Zenodo (https://doi.org/10.5281/zenodo.752340).’ The paper has been corrected online. Johannes Köster, Myles Brown, Xiaole Shirley Liu |
Bioinform. | 1 |
| 2018 | Snakemake - a scalable bioinformatics workflow engine
Johannes Köster, Sven Rahmann |
Bioinform. | 1 |
| 2018 | VIPER: Visualization Pipeline for RNA-seq, a Snakemake workflow for efficient and complete RNA-seq analysisabstractBACKGROUND: RNA sequencing has become a ubiquitous technology used throughout life sciences as an effective method of measuring RNA abundance quantitatively in tissues and cells. The increase in use of RNA-seq technology has led to the continuous development of new tools for every step of analysis from alignment to downstream pathway analysis. However, effectively using these analysis tools in a scalable and reproducible way can be challenging, especially for non-experts. RESULTS: Using the workflow management system Snakemake we have developed a user friendly, fast, efficient, and comprehensive pipeline for RNA-seq analysis. VIPER (Visualization Pipeline for RNA-seq analysis) is an analysis workflow that combines some of the most popular tools to take RNA-seq analysis from raw sequencing data, through alignment and quality control, into downstream differential expression and pathway analysis. VIPER has been created in a modular fashion to allow for the rapid incorporation of new tools to expand the capabilities. This capacity has already been exploited to include very recently developed tools that explore immune infiltrate and T-cell CDR (Complementarity-Determining Regions) reconstruction abilities. The pipeline has been conveniently packaged such that minimal computational skills are required to download and install the dozens of software packages that VIPER uses. CONCLUSIONS: VIPER is a comprehensive solution that performs most standard RNA-seq analyses quickly and effectively with a built-in capacity for customization and expansion. MacIntosh Cornwell, Mahesh Vangala, Len Taing, Zachary Herbert, Johannes Köster, Hanfei Sun, Taiwen Li, Xintao Qiu, Matthew Pun, Rinath Jeselsohn, Myles Brown, Xiaole Shirley Liu, Henry W. Long |
BMC Bioinform. | 5 |
| 2016 | Rust-Bio: a fast and safe bioinformatics libraryabstractSUMMARY: We present Rust-Bio, the first general purpose bioinformatics library for the innovative Rust programming language. Rust-Bio leverages the unique combination of speed, memory safety and high-level syntax offered by Rust to provide a fast and safe set of bioinformatics algorithms and data structures with a focus on sequence analysis. AVAILABILITY AND IMPLEMENTATION: Rust-Bio is available open source under the MIT license at https://rust-bio.github.io. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Johannes Köster |
Bioinform. | 1 |
| 2016 | CRISPR-DO for genome-wide CRISPR design and optimizationabstractMOTIVATION: Despite the growing popularity in using CRISPR/Cas9 technology for genome editing and gene knockout, its performance still relies on well-designed single guide RNAs (sgRNA). In this study, we propose a web application for the Design and Optimization (CRISPR-DO) of guide sequences that target both coding and non-coding regions in spCas9 CRISPR system across human, mouse, zebrafish, fly and worm genomes. CRISPR-DO uses a computational sequence model to predict sgRNA efficiency, and employs a specificity scoring function to evaluate the potential of off-target effect. It also provides information on functional conservation of target sequences, as well as the overlaps with exons, putative regulatory sequences and single-nucleotide polymorphisms (SNPs). The web application has a user-friendly genome-browser interface to facilitate the selection of the best target DNA sequences for experimental design. AVAILABILITY AND IMPLEMENTATION: CRISPR-DO is available at http://cistrome.org/crispr/ CONTACT: [email protected] or [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online. Johannes Köster, Sheng'en Hu, Wei Li 0035, Qingyi Cao, Jinzeng Wang, Shenglin Mei, Xiaole Shirley Liu |
Bioinform. | 2 |
| 2016 | SimLoRD: Simulation of Long Read DataabstractMOTIVATION: Third generation sequencing methods provide longer reads than second generation methods and have distinct error characteristics. While there exist many read simulators for second generation data, there is a very limited choice for third generation data. RESULTS: We analyzed public data from Pacific Biosciences (PacBio) SMRT sequencing, developed an error model and implemented it in a new read simulator called SimLoRD. It offers options to choose the read length distribution and to model error probabilities depending on the number of passes through the sequencer. The new error model makes SimLoRD the most realistic SMRT read simulator available. AVAILABILITY AND IMPLEMENTATION: SimLoRD is available open source at http://bitbucket.org/genomeinformatics/simlord/ and installable via Bioconda (http://bioconda.github.io). CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bianca K. Stöcker, Johannes Köster, Sven Rahmann |
Bioinform. | 2 |
| 2012 | Snakemake - a scalable bioinformatics workflow engineabstractSUMMARY: Snakemake is a workflow engine that provides a readable Python-based workflow definition language and a powerful execution environment that scales from single-core workstations to compute clusters without modifying the workflow. It is the first system to support the use of automatically inferred multiple named wildcards (or variables) in input and output filenames. AVAILABILITY: http://snakemake.googlecode.com. CONTACT: [email protected]. Johannes Köster, Sven Rahmann |
Bioinform. | 1 |