Pjotr Prins

dblp:01/8712 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-8021-9162ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Pangenome-Informed Language Models for Synthetic Genome Sequence Generation
abstract
Language Models (LM) have been extensively utilized for learning DNA sequence patterns and generating synthetic sequences. In this paper, we present a novel approach for the generation of synthetic DNA data using pangenomes in combination with LM. We introduce three innovative pangenome-based tokenization schemes that enhance DNA sequence generation. Our experimental results demonstrate the superiority of pangenome-based tokenization over classical methods in generating high-utility synthetic DNA sequences, highlighting significant improvements in training efficiency and sequence quality.
Pengzhi Huang, François Charton, Jan-Niklas Schmelzle, Shelby S. Darnell, Pjotr Prins, Erik Garrison, G. Edward Suh
BIBM5
2024 Rapid GPU-Based Pangenome Graph Layout
abstract
Computational Pangenomics is an emerging field that studies genetic variation using a graph structure encompassing multiple genomes. Visualizing pangenome graphs is vital for understanding genome diversity. Yet, handling large graphs can be challenging due to the high computational demands of the graph layout process. In this work, we conduct a thorough performance characterization of a state-of-the-art pangenome graph layout algorithm, revealing significant data-level parallelism, which makes GPUs a promising option for compute acceleration. However, irregular data access and the algorithm’s memory-bound nature present significant hurdles. To overcome these challenges, we develop a solution implementing three key optimizations: a cache-friendly data layout, coalesced random states, and warp merging. Additionally, we propose a quantitative metric for scalable evaluation of pangenome layout quality. Evaluated on 24 human whole-chromosome pangenomes, our GPU-based solution achieves a 57.3x speedup over the state-of-the-art multithreaded CPU baseline without layout quality loss, reducing execution time from hours to minutes.
Jiajie Li 0008, Jan-Niklas Schmelzle, Yixiao Du, Simon Heumos, Andrea Guarracino, Giulia Guidi, Pjotr Prins, Erik Garrison, Zhiru Zhang
SC7
2024 Pangenome graph layout by Path-Guided Stochastic Gradient Descent
abstract
MOTIVATION: The increasing availability of complete genomes demands for models to study genomic variability within entire populations. Pangenome graphs capture the full genomic similarity and diversity between multiple genomes. In order to understand them, we need to see them. For visualization, we need a human-readable graph layout: a graph embedding in low (e.g. two) dimensional depictions. Due to a pangenome graph's potential excessive size, this is a significant challenge. RESULTS: In response, we introduce a novel graph layout algorithm: the Path-Guided Stochastic Gradient Descent (PG-SGD). PG-SGD uses the genomes, represented in the pangenome graph as paths, as an embedded positional system to sample genomic distances between pairs of nodes. This avoids the quadratic cost seen in previous versions of graph drawing by SGD. We show that our implementation efficiently computes the low-dimensional layouts of gigabase-scale pangenome graphs, unveiling their biological features. AVAILABILITY AND IMPLEMENTATION: We integrated PG-SGD in ODGI which is released as free software under the MIT open source license. Source code is available at https://github.com/pangenome/odgi.
Simon Heumos, Andrea Guarracino, Jan-Niklas Schmelzle, Jiajie Li 0008, Zhiru Zhang, Jörg Hagmann, Sven Nahnsen, Pjotr Prins, Erik Garrison
Bioinform.8
2024 Cluster-efficient pangenome graph construction with nf-core/pangenome
abstract
MOTIVATION: Pangenome graphs offer a comprehensive way of capturing genomic variability across multiple genomes. However, current construction methods often introduce biases, excluding complex sequences or relying on references. The PanGenome Graph Builder (PGGB) addresses these issues. To date, though, there is no state-of-the-art pipeline allowing for easy deployment, efficient and dynamic use of available resources, and scalable usage at the same time. RESULTS: To overcome these limitations, we present nf-core/pangenome, a reference-unbiased approach implemented in Nextflow following nf-core's best practices. Leveraging biocontainers ensures portability and seamless deployment in High-Performance Computing (HPC) environments. Unlike PGGB, nf-core/pangenome distributes alignments across cluster nodes, enabling scalability. Demonstrating its efficiency, we constructed pangenome graphs for 1000 human chromosome 19 haplotypes and 2146 Escherichia coli sequences, achieving a two to threefold speedup compared to PGGB without increasing greenhouse gas emissions. AVAILABILITY AND IMPLEMENTATION: nf-core/pangenome is released under the MIT open-source license, available on GitHub and Zenodo, with documentation accessible at https://nf-co.re/pangenome/docs/usage.
Simon Heumos, Michael L. Heuer, Friederike Hanssen, Lukas Heumos, Andrea Guarracino, Peter Heringer, Philipp Ehmele, Pjotr Prins, Erik Garrison, Sven Nahnsen
Bioinform.8
2022 ODGI: understanding pangenome graphs
abstract
MOTIVATION: Pangenome graphs provide a complete representation of the mutual alignment of collections of genomes. These models offer the opportunity to study the entire genomic diversity of a population, including structurally complex regions. Nevertheless, analyzing hundreds of gigabase-scale genomes using pangenome graphs is difficult as it is not well-supported by existing tools. Hence, fast and versatile software is required to ask advanced questions to such data in an efficient way. RESULTS: We wrote Optimized Dynamic Genome/Graph Implementation (ODGI), a novel suite of tools that implements scalable algorithms and has an efficient in-memory representation of DNA pangenome graphs in the form of variation graphs. ODGI supports pre-built graphs in the Graphical Fragment Assembly format. ODGI includes tools for detecting complex regions, extracting pangenomic loci, removing artifacts, exploratory analysis, manipulation, validation and visualization. Its fast parallel execution facilitates routine pangenomic tasks, as well as pipelines that can quickly answer complex biological questions of gigabase-scale pangenome graphs. AVAILABILITY AND IMPLEMENTATION: ODGI is published as free software under the MIT open source license. Source code can be downloaded from https://github.com/pangenome/odgi and documentation is available at https://odgi.readthedocs.io. ODGI can be installed via Bioconda https://bioconda.github.io/recipes/odgi/README.html or GNU Guix https://github.com/pangenome/odgi/blob/master/guix.scm. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrea Guarracino, Simon Heumos, Sven Nahnsen, Pjotr Prins, Erik Garrison
Bioinform.4
2022 A spectrum of free software tools for processing the VCF variant call format: vcflib, bio-vcf, cyvcf2, hts-nim and slivar
abstract
Since its introduction in 2011 the variant call format (VCF) has been widely adopted for processing DNA and RNA variants in practically all population studies-as well as in somatic and germline mutation studies. The VCF format can represent single nucleotide variants, multi-nucleotide variants, insertions and deletions, and simple structural variants called and anchored against a reference genome. Here we present a spectrum of over 125 useful, complimentary free and open source software tools and libraries, we wrote and made available through the multiple vcflib, bio-vcf, cyvcf2, hts-nim and slivar projects. These tools are applied for comparison, filtering, normalisation, smoothing and annotation of VCF, as well as output of statistics, visualisation, and transformations of files variants. These tools run everyday in critical biomedical pipelines and countless shell scripts. Our tools are part of the wider bioinformatics ecosystem and we highlight best practices. We shortly discuss the design of VCF, lessons learnt, and how we can address more complex variation through pangenome graph formats, variation that can not easily be represented by the VCF format.
Erik Garrison, Zev N. Kronenberg, Eric T. Dawson, Brent S. Pedersen, Pjotr Prins
PLoS Comput. Biol.5
2020 Ten simple rules to run a successful BioHackathon
abstract
Scientific conferences are one of the most common venues for researchers and others to present and exchange new findings.In recent years, "unconferences"-i.e., meetings that promote spontaneous discussions rather than predetermined presentations-have emerged as a more collaborative approach.Unconferences differ from conferences in key ways: while conferences have a predefined set of presenters together with an audience, unconferences promote more collaborative and spontaneous interactions across participants [1].A "hackathon" is a special kind of unconference in which people come together to state, discuss, and solve problems by means of collaborative brainstorming, modeling, design, coding, testing, and documenting [2].Despite the "hack" portion in the term, hackathons welcome not only software developers but anyone involved in creating solutions that can be later consumed or exposed via software.In addition to collaboration, hackathons promote community development around a subject used as the main hackathon topic.Such a topic can correspond to a knowledge domain (e.g., semantics or genomics) or be related to a particular organization (e.g., data and data services offered).There are hackathon-like events that are called by other names, such as "codefests."In this paper, we include such events under the umbrella term "hackathon."Hackathons are especially useful in bringing together interdisciplinary sets of domain experts and specialized computer scientists with various degrees of experience and skills to "hack" solutions related to scientific topics of mutual interest.While traditional conferences focus more on transferring knowledge, hackathons are more about collaboratively generating solutions.The interactions often lead to productive collaborations, professional development opportunities, and a network of resources.Well-run hackathons are very effective at building a community among participants.Hackathons provide an opportunity for researchers and developers to interact and brainstorm with other participants in an appropriate environment to accelerate collaborations on topics of mutual benefit.In addition, they can provide a unique opportunity to think through a problem, without the usual distractions.Therefore, hackathons can be very productive and result in a major impact on the targeted topic and/or community [3].However, organizing a successful large-scale hackathon takes significant time and effort; e.g., organizing committees for the National Bioscience Database Center (NBDC) [4]/Database Center for Life Science (DBCLS) [5] and the ELIXIR Europe BioHackathons start approximately 1 year in advance.We use the term BioHackathon to refer to those hackathons addressing problems in domains related to biomedical and life sciences.BioHackathons are recognized as having a
Leyla Jael Castro, Erick Antezana, Alexander García Castro, Evan Bolton, Rafael C. Jiménez, Pjotr Prins, Juan M. Banda, Toshiaki Katayama
PLoS Comput. Biol.6
2017 Robust Cross-Platform Workflows: How Technical and Scientific Communities Collaborate to Develop, Test and Share Best Practices for Data Analysis
abstract
Information integration and workflow technologies for data analysis have always been major fields of investigation in bioinformatics. A range of popular workflow suites are available to support analyses in computational biology. Commercial providers tend to offer prepared applications remote to their clients. However, for most academic environments with local expertise, novel data collection techniques or novel data analysis, it is essential to have all the flexibility of open-source tools and open-source workflow descriptions. Workflows in data-driven science such as computational biology have considerably gained in complexity. New tools or new releases with additional features arrive at an enormous pace, and new reference data or concepts for quality control are emerging. A well-abstracted workflow and the exchange of the same across work groups have an enormous impact on the efficiency of research and the further development of the field. High-throughput sequencing adds to the avalanche of data available in the field; efficient computation and, in particular, parallel execution motivate the transition from traditional scripts and Makefiles to workflows. We here review the extant software development and distribution model with a focus on the role of integration testing and discuss the effect of common workflow language on distributions of open-source scientific software to swiftly and reliably provide the tools demanded for the execution of such formally described workflows. It is contended that, alleviated from technical differences for the execution on local machines, clusters or the cloud, communities also gain the technical means to test workflow-driven interaction across several software packages.
Steffen Möller, Stuart W. Prescott, Lars Wirzenius, Petter Reinholdtsen, Brad A. Chapman, Pjotr Prins, Stian Soiland-Reyes, Fabian Klötzl, Andrea Bagnacani, Matús Kalas, Andreas Tille, Michael R. Crusoe
Data Sci. Eng.6
2015 Sambamba: fast processing of NGS alignment formats
abstract
UNLABELLED: Sambamba is a high-performance robust tool and library for working with SAM, BAM and CRAM sequence alignment files; the most common file formats for aligned next generation sequencing data. Sambamba is a faster alternative to samtools that exploits multi-core processing and dramatically reduces processing time. Sambamba is being adopted at sequencing centers, not only because of its speed, but also because of additional functionality, including coverage analysis and powerful filtering capability. AVAILABILITY AND IMPLEMENTATION: Sambamba is free and open source software, available under a GPLv2 license. Sambamba can be downloaded and installed from http://www.open-bio.org/wiki/Sambamba.Sambamba v0.5.0 was released with doi:10.5281/zenodo.13200.
Artem Tarasov, Albert J. Vilella, Edwin Cuppen, Isaac J. Nijman, Pjotr Prins
Bioinform.5
2014 Community-driven development for computational biology at Sprints, Hackathons and Codefests
abstract
BACKGROUND: Computational biology comprises a wide range of technologies and approaches. Multiple technologies can be combined to create more powerful workflows if the individuals contributing the data or providing tools for its interpretation can find mutual understanding and consensus. Much conversation and joint investigation are required in order to identify and implement the best approaches. Traditionally, scientific conferences feature talks presenting novel technologies or insights, followed up by informal discussions during coffee breaks. In multi-institution collaborations, in order to reach agreement on implementation details or to transfer deeper insights in a technology and practical skills, a representative of one group typically visits the other. However, this does not scale well when the number of technologies or research groups is large. Conferences have responded to this issue by introducing Birds-of-a-Feather (BoF) sessions, which offer an opportunity for individuals with common interests to intensify their interaction. However, parallel BoF sessions often make it hard for participants to join multiple BoFs and find common ground between the different technologies, and BoFs are generally too short to allow time for participants to program together. RESULTS: This report summarises our experience with computational biology Codefests, Hackathons and Sprints, which are interactive developer meetings. They are structured to reduce the limitations of traditional scientific meetings described above by strengthening the interaction among peers and letting the participants determine the schedule and topics. These meetings are commonly run as loosely scheduled "unconferences" (self-organized identification of participants and topics for meetings) over at least two days, with early introductory talks to welcome and organize contributors, followed by intensive collaborative coding sessions. We summarise some prominent achievements of those meetings and describe differences in how these are organised, how their audience is addressed, and their outreach to their respective communities. CONCLUSIONS: Hackathons, Codefests and Sprints share a stimulating atmosphere that encourages participants to jointly brainstorm and tackle problems of shared interest in a self-driven proactive environment, as well as providing an opportunity for new participants to get involved in collaborative projects.
Steffen Möller, Enis Afgan, Michael Banck, Raoul Jean Pierre Bonnal, Tim Booth, John Chilton, Peter J. A. Cock, Markus Gumbel, Nomi L. Harris, Richard C. G. Holland, Matús Kalas, László Kaján, Eri Kibukawa, David R. Powell, Pjotr Prins, Jacqueline Quinn, Olivier Sallou, Francesco Strozzi, Torsten Seemann, Clare Sloggett, Stian Soiland-Reyes, William Spooner, Sascha Steinbiss, Andreas Tille, Anthony J. Travis, Roman Guimera, Toshiaki Katayama, Brad A. Chapman
BMC Bioinform.15
2012 High-throughput Molecular Docking Now in Reach for a Wider Biochemical Community
abstract
In silico molecular docking is used to predict how a small molecule, the ligand, interacts with a target protein, its receptor. Together with experimental methods like NMR or X-ray crystallography, industrial and academic groups use it for their investigation of compounds with the potential to modulate the protein's function and become a lead molecule for drug development. The interpretation of raw data, from NMR, mass spectrometry or crystallography, is greatly assisted by computers. Biochemists can, perhaps better than anybody else, perform individual analyses for the computational modeling of interactions. However, an extension towards virtual screening of compound libraries, i.e. the computational docking of thousands or even millions of ligands to a target receptor, is often perceived as technically challenging and computationally too expensive in the biochemical community. Here we describe how to integrate spare resources of regular desktop computers on and off campus using the Berkeley Open Infrastructure for Network Computing (BOINC) infrastructure for volunteer grid computing. We have brought both the BOINC server and the Auto Dock software suite into the Debian and Ubuntu Linux distributions, and provide detailed instructions to help render the implementation of a large-scale high-throughput docking (HTD) project straightforward. Thus, this increased availability of computational resources, protocols and source code opens up the possibility of many new self-run projects and collaborations, for those who may be adopting the technology for the first time. On the technical level, we expect to observe more contributors from the Open Source and academic communities. On the biological side, we anticipate new and faster progress on commercially less interesting and so-called 'neglected' diseases.
Dhananjay M. Balan, Tomas Malinauskas, Pjotr Prins, Steffen Möller
PDP3
2012 Bioinformatics tools and database resources for systems genetics analysis in mice - a short review and an evaluation of future needs
abstract
During a meeting of the SYSGENET working group 'Bioinformatics', currently available software tools and databases for systems genetics in mice were reviewed and the needs for future developments discussed. The group evaluated interoperability and performed initial feasibility studies. To aid future compatibility of software and exchange of already developed software modules, a strong recommendation was made by the group to integrate HAPPY and R/qtl analysis toolboxes, GeneNetwork and XGAP database platforms, and TIQS and xQTL processing platforms. R should be used as the principal computer language for QTL data analysis in all platforms and a 'cloud' should be used for software dissemination to the community. Furthermore, the working group recommended that all data models and software source code should be made visible in public repositories to allow a coordinated effort on the use of common data structures and file formats.
Caroline Durrant, Morris A. Swertz, Rudi Alberts, Danny Arends, Steffen Möller, Richard Mott, Pjotr Prins, K. Joeri van der Velde, Ritsert C. Jansen, Klaus Schughart
Briefings Bioinform.7
2012 xQTL workbench: a scalable web environment for multi-level QTL analysis
abstract
SUMMARY: xQTL workbench is a scalable web platform for the mapping of quantitative trait loci (QTLs) at multiple levels: for example gene expression (eQTL), protein abundance (pQTL), metabolite abundance (mQTL) and phenotype (phQTL) data. Popular QTL mapping methods for model organism and human populations are accessible via the web user interface. Large calculations scale easily on to multi-core computers, clusters and Cloud. All data involved can be uploaded and queried online: markers, genotypes, microarrays, NGS, LC-MS, GC-MS, NMR, etc. When new data types come available, xQTL workbench is quickly customized using the Molgenis software generator. AVAILABILITY: xQTL workbench runs on all common platforms, including Linux, Mac OS X and Windows. An online demo system, installation guide, tutorials, software and source code are available under the LGPL3 license from http://www.xqtl.org. CONTACT: [email protected].
Danny Arends, K. Joeri van der Velde, Pjotr Prins, Karl W. Broman, Steffen Möller, Ritsert C. Jansen, Morris A. Swertz
Bioinform.3
2012 Biogem: an effective tool-based approach for scaling up open source software development in bioinformatics
abstract
SUMMARY: Biogem provides a software development environment for the Ruby programming language, which encourages community-based software development for bioinformatics while lowering the barrier to entry and encouraging best practices. Biogem, with its targeted modular and decentralized approach, software generator, tools and tight web integration, is an improved general model for scaling up collaborative open source software development in bioinformatics. AVAILABILITY: Biogem and modules are free and are OSS. Biogem runs on all systems that support recent versions of Ruby, including Linux, Mac OS X and Windows. Further information at http://www.biogems.info. A tutorial is available at http://www.biogems.info/howto.html CONTACT: [email protected].
Raoul Jean Pierre Bonnal, Jan Aerts, George Githinji, Naohisa Goto, Daniel MacLean, Chase A. Miller, Hiroyuki Mishima, Massimiliano Pagani, Ricardo Ramirez-Gonzalez, Geert Smant, Francesco Strozzi, Rob Syme, Rutger A. Vos, Trevor J. Wennblom, Ben J. Woodcroft, Toshiaki Katayama, Pjotr Prins
Bioinform.17
2010 R/qtl: high-throughput multiple QTL mapping
abstract
MOTIVATION: R/qtl is free and powerful software for mapping and exploring quantitative trait loci (QTL). R/qtl provides a fully comprehensive range of methods for a wide range of experimental cross types. We recently added multiple QTL mapping (MQM) to R/qtl. MQM adds higher statistical power to detect and disentangle the effects of multiple linked and unlinked QTL compared with many other methods. MQM for R/qtl adds many new features including improved handling of missing data, analysis of 10,000 s of molecular traits, permutation for determining significance thresholds for QTL and QTL hot spots, and visualizations for cis-trans and QTL interaction effects. MQM for R/qtl is the first free and open source implementation of MQM that is multi-platform, scalable and suitable for automated procedures and large genetical genomics datasets. AVAILABILITY: R/qtl is free and open source multi-platform software for the statistical language R, and is made available under the GPLv3 license. R/qtl can be installed from http://www.rqtl.org/. R/qtl queries should be directed at the mailing list, see http://www.rqtl.org/list/. CONTACT: [email protected].
Danny Arends, Pjotr Prins, Ritsert C. Jansen, Karl W. Broman
Bioinform.2
2010 BioRuby: bioinformatics software for the Ruby programming language
abstract
SUMMARY: The BioRuby software toolkit contains a comprehensive set of free development tools and libraries for bioinformatics and molecular biology, written in the Ruby programming language. BioRuby has components for sequence analysis, pathway analysis, protein modelling and phylogenetic analysis; it supports many widely used data formats and provides easy access to databases, external programs and public web services, including BLAST, KEGG, GenBank, MEDLINE and GO. BioRuby comes with a tutorial, documentation and an interactive environment, which can be used in the shell, and in the web browser. AVAILABILITY: BioRuby is free and open source software, made available under the Ruby license. BioRuby runs on all platforms that support Ruby, including Linux, Mac OS X and Windows. And, with JRuby, BioRuby runs on the Java Virtual Machine. The source code is available from http://www.bioruby.org/. CONTACT: [email protected]
Naohisa Goto, Pjotr Prins, Mitsuteru Nakao, Raoul Jean Pierre Bonnal, Jan Aerts, Toshiaki Katayama
Bioinform.2