J. Gregory Caporaso

dblp:73/1948 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-8865-1670ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Assessing microbiome engraftment extent following fecal microbiota transplant with q2-fmt
abstract
We present q2-fmt, a QIIME 2 plugin that provides diverse methods for assessing the extent of microbiome engraftment following fecal microbiota transplant. The methods implemented here were informed by a recent literature review on approaches for assessing FMT engraftment, and cover aspects of engraftment including Community Coalescence, Indicator Features, and Resilience. q2-fmt is free for all use, and detailed documentation illustrating worked examples on a real-world data set are provided in the project's documentation.
Chloe Herman, Evan Bolyen, Anthony Simard, Elizabeth Gehret, J. Gregory Caporaso
PLoS Comput. Biol.5
2024 Best practices to evaluate the impact of biomedical research software - metric collection beyond citations
abstract
MOTIVATION: Software is vital for the advancement of biology and medicine. Impact evaluations of scientific software have primarily emphasized traditional citation metrics of associated papers, despite these metrics inadequately capturing the dynamic picture of impact and despite challenges with improper citation. RESULTS: To understand how software developers evaluate their tools, we conducted a survey of participants in the Informatics Technology for Cancer Research (ITCR) program funded by the National Cancer Institute (NCI). We found that although developers realize the value of more extensive metric collection, they find a lack of funding and time hindering. We also investigated software among this community for how often infrastructure that supports more nontraditional metrics were implemented and how this impacted rates of papers describing usage of the software. We found that infrastructure such as social media presence, more in-depth documentation, the presence of software health metrics, and clear information on how to contact developers seemed to be associated with increased mention rates. Analysing more diverse metrics can enable developers to better understand user engagement, justify continued funding, identify novel use cases, pinpoint improvement areas, and ultimately amplify their software's impact. Challenges are associated, including distorted or misleading metrics, as well as ethical and security concerns. More attention to nuances involved in capturing impact across the spectrum of biomedical software is needed. For funders and developers, we outline guidance based on experience from our community. By considering how we evaluate software, we can empower developers to create tools that more effectively accelerate biological and medical research progress. AVAILABILITY AND IMPLEMENTATION: More information about the analysis, as well as access to data and code is available at https://github.com/fhdsl/ITCR_Metrics_manuscript_website.
Awan Afiaz, John Chamberlin, David Hanauer, Candace Savonen, Mary J. Goldman, Martin Morgan, Michael Reich, Alexander Getka, Aaron Holmes, Sarthak Pati, Dan Knight, Paul C. Boutros, Spyridon Bakas, J. Gregory Caporaso, Guilherme Del Fiol, Harry Hochheiser, Brian Haas, Patrick D. Schloss, James A. Eddy, Jake Albrecht, Andriy Fedorov, Levi Waldron, Ava M. Hoffman, Richard L. Bradshaw, Jeffrey T. Leek, Carrie Wright
Bioinform.15
2023 Facilitating bioinformatics reproducibility with QIIME 2 Provenance Replay
abstract
Study reproducibility is essential to corroborate, build on, and learn from the results of scientific research but is notoriously challenging in bioinformatics, which often involves large data sets and complex analytic workflows involving many different tools. Additionally, many biologists are not trained in how to effectively record their bioinformatics analysis steps to ensure reproducibility, so critical information is often missing. Software tools used in bioinformatics can automate provenance tracking of the results they generate, removing most barriers to bioinformatics reproducibility. Here we present an implementation of that idea, Provenance Replay, a tool for generating new executable code from results generated with the QIIME 2 bioinformatics platform, and discuss considerations for bioinformatics developers who wish to implement similar functionality in their software.
Christopher R. Keefe, Matthew R. Dillon, Elizabeth Gehret, Chloe Herman, Mary Jewell, Colin V. Wood, Evan Bolyen, J. Gregory Caporaso
PLoS Comput. Biol.8
2022 Multi-omics data integration reveals metabolome as the top predictor of the cervicovaginal microenvironment
abstract
Emerging evidence suggests that host-microbe interaction in the cervicovaginal microenvironment contributes to cervical carcinogenesis, yet dissecting these complex interactions is challenging. Herein, we performed an integrated analysis of multiple "omics" datasets to develop predictive models of the cervicovaginal microenvironment and identify characteristic features of vaginal microbiome, genital inflammation and disease status. Microbiomes, vaginal pH, immunoproteomes and metabolomes were measured in cervicovaginal specimens collected from a cohort (n = 72) of Arizonan women with or without cervical neoplasm. Multi-omics integration methods, including neural networks (mmvec) and Random Forest supervised learning, were utilized to explore potential interactions and develop predictive models. Our integrated analyses revealed that immune and cancer biomarker concentrations were reliably predicted by Random Forest regressors trained on microbial and metabolic features, suggesting close correspondence between the vaginal microbiome, metabolome, and genital inflammation involved in cervical carcinogenesis. Furthermore, we show that features of the microbiome and host microenvironment, including metabolites, microbial taxa, and immune biomarkers are predictive of genital inflammation status, but only weakly to moderately predictive of cervical neoplastic disease status. Different feature classes were important for prediction of different phenotypes. Lipids (e.g. sphingolipids and long-chain unsaturated fatty acids) were strong predictors of genital inflammation, whereas predictions of vaginal microbiota and vaginal pH relied mostly on alterations in amino acid metabolism. Finally, we identified key immune biomarkers associated with the vaginal microbiota composition and vaginal pH (MIF), as well as genital inflammation (IL-6, IL-10, MIP-1α).
Nicholas A. Bokulich, Pawel Laniewski, Anja Adamov, Dana M. Chase, J. Gregory Caporaso, Melissa M. Herbst-Kralovetz
PLoS Comput. Biol.5
2021 Experiences and lessons learned from two virtual, hands-on microbiome bioinformatics workshops
abstract
In October of 2020, in response to the Coronavirus Disease 2019 (COVID-19) pandemic, our team hosted our first fully online workshop teaching the QIIME 2 microbiome bioinformatics platform. We had 75 enrolled participants who joined from at least 25 different countries on 6 continents, and we had 22 instructors on 4 continents. In the 5-day workshop, participants worked hands-on with a cloud-based shared compute cluster that we deployed for this course. The event was well received, and participants provided feedback and suggestions in a postworkshop questionnaire. In January of 2021, we followed this workshop with a second fully online workshop, incorporating lessons from the first. Here, we present details on the technology and protocols that we used to run these workshops, focusing on the first workshop and then introducing changes made for the second workshop. We discuss what worked well, what didn't work well, and what we plan to do differently in future workshops.
Matthew R. Dillon, Evan Bolyen, Anja Adamov, Aeriel Belk, Emily Borsom, Zachary Burcham, Justine W. Debelius, Heather Deel, Alex Emmons, Mehrbod Estaki, Chloe Herman, Christopher R. Keefe, Jamie T. Morton, Renato R. M. Oliveira, Andrew Sanchez, Anthony Simard, Yoshiki Vazquez-Baeza, Michal Ziemski, Hazuki E. Miwa, Terry A. Kerere, Carline Coote, Richard Bonneau, Rob Knight 0001, Guilherme C. Oliveira 0001, Piraveen Gopalasingam, Benjamin D. Kaehler, Emily K. Cope, Jessica L. Metcalf, Michael S. Robeson II, Nicholas A. Bokulich, J. Gregory Caporaso
PLoS Comput. Biol.31
2011 TopiaryExplorer: visualizing large phylogenetic trees with environmental metadata
abstract
MOTIVATION: Microbial community profiling is a highly active area of research, but tools that facilitate visualization of phylogenetic trees and associated environmental data have not kept up with the increasing quantity of data generated in these studies. RESULTS: TopiaryExplorer supports the visualization of very large phylogenetic trees, including features such as the automated coloring of branches by environmental data, manipulation of trees and incorporation of per-tip metadata (e.g. taxonomic labels). AVAILABILITY: http://topiaryexplorer.sourceforge.net. CONTACT: [email protected].
Meg Pirrung, Ryan Kennedy, J. Gregory Caporaso, Jesse Stombaugh, Doug Wendel, Rob Knight 0001
Bioinform.3
2011 PrimerProspector: de novo design and taxonomic analysis of barcoded polymerase chain reaction primers
abstract
MOTIVATION: PCR amplification of DNA is a key preliminary step in many applications of high-throughput sequencing technologies, yet design of novel barcoded primers and taxonomic analysis of novel or existing primers remains a challenging task. RESULTS: PrimerProspector is an open-source software package that allows researchers to develop new primers from collections of sequences and to evaluate existing primers in the context of taxonomic data. AVAILABILITY: PrimerProspector is open-source software available at http://pprospector.sourceforge.net CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
William A. Walters, J. Gregory Caporaso, Christian L. Lauber, Donna Berg-Lyons, Noah Fierer, Rob Knight 0001
Bioinform.2
2010 PyNAST: a flexible tool for aligning sequences to a template alignment
abstract
MOTIVATION: The Nearest Alignment Space Termination (NAST) tool is commonly used in sequence-based microbial ecology community analysis, but due to the limited portability of the original implementation, it has not been as widely adopted as possible. Python Nearest Alignment Space Termination (PyNAST) is a complete reimplementation of NAST, which includes three convenient interfaces: a Mac OS X GUI, a command-line interface and a simple application programming interface (API). RESULTS: The availability of PyNAST will make the popular NAST algorithm more portable and thereby applicable to datasets orders of magnitude larger by allowing users to install PyNAST on their own hardware. Additionally because users can align to arbitrary template alignments, a feature not available via the original NAST web interface, the NAST algorithm will be readily applicable to novel tasks outside of microbial community analysis. AVAILABILITY: PyNAST is available at http://pynast.sourceforge.net.
J. Gregory Caporaso, Kyle Bittinger, Frederic D. Bushman, Todd Z. DeSantis, Gary L. Andersen, Rob Knight 0001
Bioinform.1
2007 MutationFinder: a high-performance system for extracting point mutation mentions from text
abstract
Discussion of point mutations is ubiquitous in biomedical literature, and manually compiling databases or literature on mutations in specific genes or proteins is tedious. We present an open-source, rule-based system, MutationFinder, for extracting point mutation mentions from text. On blind test data, it achieves nearly perfect precision and a markedly improved recall over a baseline. AVAILABILITY: MutationFinder, along with a high-quality gold standard data set, and a scoring script for mutation extraction systems have been made publicly available. Implementations, source code and unit tests are available in Python, Perl and Java. MutationFinder can be used as a stand-alone script, or imported by other applications. PROJECT URL: http://bionlp.sourceforge.net.
J. Gregory Caporaso, William A. Baumgartner Jr., David A. Randolph, Kevin Cohen 0001, Lawrence Hunter
Bioinform.1