VLDB 2026 Research / reviewers in the wild / expert
Nicholas A. Bokulich
dblp:229/4861
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0002-1784-8935ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ViromeXplore: integrative workflows for complete and reproducible virome characterizationabstractViruses play a crucial role in shaping microbial communities and global biogeochemical cycles, yet their vast genetic diversity remains underexplored. Next-generation sequencing technologies allow untargeted profiling of metagenomes from viral communities (viromes). However, existing workflows often lack modularity, flexibility, and seamless integration with other microbiome analysis platforms. Here, we introduce "ViromeXplore," a set of modular Nextflow workflows designed for efficient virome analysis. ViromeXplore incorporates state-of-the-art tools for contamination estimation, viral sequence identification, taxonomic assignment, functional annotation, and host prediction while optimizing computational resources. The workflows are containerized using Docker and Singularity, ensuring reproducibility and ease of deployment. Additionally, ViromeXplore offers optional integration with QIIME 2 and MOSHPIT, facilitating provenance tracking and interoperability with microbiome bioinformatics pipelines. By providing a scalable, user-friendly, and computationally efficient framework, ViromeXplore enhances viral metagenomic analysis and contributes to a deeper understanding of viral ecology. ViromeXplore is freely available at https://github.com/rhernandvel/ViromeXplore. Rodrigo Hernández-Velázquez, Michal Ziemski, Nicholas A. Bokulich |
Briefings Bioinform. | 3 |
| 2022 | Reproducible acquisition, management and meta-analysis of nucleotide sequence (meta)data using q2-fondueabstractMOTIVATION: The volume of public nucleotide sequence data has blossomed over the past two decades and is ripe for re- and meta-analyses to enable novel discoveries. However, reproducible re-use and management of sequence datasets and associated metadata remain critical challenges. We created the open source Python package q2-fondue to enable user-friendly acquisition, re-use and management of public sequence (meta)data while adhering to open data principles. RESULTS: q2-fondue allows fully provenance-tracked programmatic access to and management of data from the NCBI Sequence Read Archive (SRA). Unlike other packages allowing download of sequence data from the SRA, q2-fondue enables full data provenance tracking from data download to final visualization, integrates with the QIIME 2 ecosystem, prevents data loss upon space exhaustion and allows download of (meta)data given a publication library. To highlight its manifold capabilities, we present executable demonstrations using publicly available amplicon, whole genome and metagenome datasets. AVAILABILITY AND IMPLEMENTATION: q2-fondue is available as an open-source BSD-3-licensed Python package at https://github.com/bokulich-lab/q2-fondue. Usage tutorials are available in the same repository. All Jupyter notebooks used in this article are available under https://github.com/bokulich-lab/q2-fondue-examples. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michal Ziemski, Anja Adamov, Lina Kim, Lena Flörl, Nicholas A. Bokulich |
Bioinform. | 5 |
| 2022 | Multi-omics data integration reveals metabolome as the top predictor of the cervicovaginal microenvironmentabstractEmerging evidence suggests that host-microbe interaction in the cervicovaginal microenvironment contributes to cervical carcinogenesis, yet dissecting these complex interactions is challenging. Herein, we performed an integrated analysis of multiple "omics" datasets to develop predictive models of the cervicovaginal microenvironment and identify characteristic features of vaginal microbiome, genital inflammation and disease status. Microbiomes, vaginal pH, immunoproteomes and metabolomes were measured in cervicovaginal specimens collected from a cohort (n = 72) of Arizonan women with or without cervical neoplasm. Multi-omics integration methods, including neural networks (mmvec) and Random Forest supervised learning, were utilized to explore potential interactions and develop predictive models. Our integrated analyses revealed that immune and cancer biomarker concentrations were reliably predicted by Random Forest regressors trained on microbial and metabolic features, suggesting close correspondence between the vaginal microbiome, metabolome, and genital inflammation involved in cervical carcinogenesis. Furthermore, we show that features of the microbiome and host microenvironment, including metabolites, microbial taxa, and immune biomarkers are predictive of genital inflammation status, but only weakly to moderately predictive of cervical neoplastic disease status. Different feature classes were important for prediction of different phenotypes. Lipids (e.g. sphingolipids and long-chain unsaturated fatty acids) were strong predictors of genital inflammation, whereas predictions of vaginal microbiota and vaginal pH relied mostly on alterations in amino acid metabolism. Finally, we identified key immune biomarkers associated with the vaginal microbiota composition and vaginal pH (MIF), as well as genital inflammation (IL-6, IL-10, MIP-1α). Nicholas A. Bokulich, Pawel Laniewski, Anja Adamov, Dana M. Chase, J. Gregory Caporaso, Melissa M. Herbst-Kralovetz |
PLoS Comput. Biol. | 1 |
| 2021 | Experiences and lessons learned from two virtual, hands-on microbiome bioinformatics workshopsabstractIn October of 2020, in response to the Coronavirus Disease 2019 (COVID-19) pandemic, our team hosted our first fully online workshop teaching the QIIME 2 microbiome bioinformatics platform. We had 75 enrolled participants who joined from at least 25 different countries on 6 continents, and we had 22 instructors on 4 continents. In the 5-day workshop, participants worked hands-on with a cloud-based shared compute cluster that we deployed for this course. The event was well received, and participants provided feedback and suggestions in a postworkshop questionnaire. In January of 2021, we followed this workshop with a second fully online workshop, incorporating lessons from the first. Here, we present details on the technology and protocols that we used to run these workshops, focusing on the first workshop and then introducing changes made for the second workshop. We discuss what worked well, what didn't work well, and what we plan to do differently in future workshops. Matthew R. Dillon, Evan Bolyen, Anja Adamov, Aeriel Belk, Emily Borsom, Zachary Burcham, Justine W. Debelius, Heather Deel, Alex Emmons, Mehrbod Estaki, Chloe Herman, Christopher R. Keefe, Jamie T. Morton, Renato R. M. Oliveira, Andrew Sanchez, Anthony Simard, Yoshiki Vazquez-Baeza, Michal Ziemski, Hazuki E. Miwa, Terry A. Kerere, Carline Coote, Richard Bonneau, Rob Knight 0001, Guilherme C. Oliveira 0001, Piraveen Gopalasingam, Benjamin D. Kaehler, Emily K. Cope, Jessica L. Metcalf, Michael S. Robeson II, Nicholas A. Bokulich, J. Gregory Caporaso |
PLoS Comput. Biol. | 30 |
| 2021 | RESCRIPt: Reproducible sequence taxonomy reference database managementabstractNucleotide sequence and taxonomy reference databases are critical resources for widespread applications including marker-gene and metagenome sequencing for microbiome analysis, diet metabarcoding, and environmental DNA (eDNA) surveys. Reproducibly generating, managing, using, and evaluating nucleotide sequence and taxonomy reference databases creates a significant bottleneck for researchers aiming to generate custom sequence databases. Furthermore, database composition drastically influences results, and lack of standardization limits cross-study comparisons. To address these challenges, we developed RESCRIPt, a Python 3 software package and QIIME 2 plugin for reproducible generation and management of reference sequence taxonomy databases, including dedicated functions that streamline creating databases from popular sources, and functions for evaluating, comparing, and interactively exploring qualitative and quantitative characteristics across reference databases. To highlight the breadth and capabilities of RESCRIPt, we provide several examples for working with popular databases for microbiome profiling (SILVA, Greengenes, NCBI-RefSeq, GTDB), eDNA and diet metabarcoding surveys (BOLD, GenBank), as well as for genome comparison. We show that bigger is not always better, and reference databases with standardized taxonomies and those that focus on type strains have quantitative advantages, though may not be appropriate for all use cases. Most databases appear to benefit from some curation (quality filtering), though sequence clustering appears detrimental to database quality. Finally, we demonstrate the breadth and extensibility of RESCRIPt for reproducible workflows with a comparison of global hepatitis genomes. RESCRIPt provides tools to democratize the process of reference database acquisition and management, enabling researchers to reproducibly and transparently create reference materials for diverse research applications. RESCRIPt is released under a permissive BSD-3 license at https://github.com/bokulich-lab/RESCRIPt. Michael S. Robeson II, Devon R. O'Rourke, Benjamin D. Kaehler, Michal Ziemski, Matthew R. Dillon, Jeffrey T. Foster, Nicholas A. Bokulich |
PLoS Comput. Biol. | 7 |