Joseph N. Paulson

dblp:181/6714 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
4since 2021 · last 2023
0000-0001-8221-7139ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › computational microbiology
microbiome analysis
0.932023
MicrobiomeExplorer: an R package for the analysis and visualization of microbial communities · Bioinform. 2021
Privacy-preserving microbiome analysis using secure computation · Bioinform. 2016
mbQTL: an R/Bioconductor package for microbial quantitative trait loci (QTL) estimation · Bioinform. 2023
Bioinformatics and computational biology
metagenomics
0.622019
metagenomeFeatures: an R package for working with 16S rRNA reference databases and marker-gene survey feature data · Bioinform. 2019
Privacy-preserving microbiome analysis using secure computation · Bioinform. 2016
Bioinformatics and computational biology
feature selection
0.612022
GameRank: R package for feature selection and construction · Bioinform. 2022
Bioinformatics and computational biology › computational microbiology › microbiome analysis
microbial community analysis
0.512021
MicrobiomeExplorer: an R package for the analysis and visualization of microbial communities · Bioinform. 2021
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.312017
Estimating gene regulatory networks with pandaR · Bioinform. 2017
Privacy and data protection
privacy-preserving data analysis
0.212016
Privacy-preserving microbiome analysis using secure computation · Bioinform. 2016

Methods — techniques the papers use, named apart from their topics

statistical association testing · 0.7maximum likelihood · 0.6combinatorial search · 0.6shiny application · 0.5secure multiparty computation · 0.5r package · 0.5network data assimilation · 0.3message passing · 0.3
YearPublicationVenuePosition
2023 mbQTL: an R/Bioconductor package for microbial quantitative trait loci (QTL) estimation
abstract
MOTIVATION: In recent years, significant strides have been made in the field of genomics, with the commencement of large-scale studies aimed at collecting host mutational profiles and microbiome data. The amalgamation of host gene mutational profiles in both healthy and diseased subjects with microbial abundance data holds immense promise in providing insights into several crucial research questions, including the development and progression of diseases, as well as individual responses to therapeutic interventions. With the advent of sequencing methods such as 16s ribosomal RNA (rRNA) sequencing and whole genome sequencing, there is increasing evidence of interplay of human genetics and microbial communities. Quantitative trait loci associated with microbial abundance (mbQTLs), are genetic variants that influence the abundance of microbial populations within the host. RESULTS: Here, we introduce mbQTL, the first R package integrating 16S ribosomal RNA (rRNA) sequencing and single-nucleotide variation (SNV) and single-nucleotide polymorphism (SNP) data. We describe various statistical methods implemented for the identification of microbe-SNV pairs, relevant statistical measures, and plot functionality for interpretation. AVAILABILITY AND IMPLEMENTATION: mbQTL is available on bioconductor at https://bioconductor.org/packages/mbQTL/.
Mercedeh Movassagh, Steven J. Schiff, Joseph N. Paulson
Bioinform.3
2022 GameRank: R package for feature selection and construction
abstract
MOTIVATION: Building calibrated and discriminating predictive models can be developed through the direct optimization of model performance metrics with combinatorial search algorithms. Often, predictive algorithms are desired in clinical settings to identify patients that may be high and low risk. However, due to the large combinatorial search space, these algorithms are slow and do not guarantee the global optimality of their selection. RESULTS: Here, we present a novel and quick maximum likelihood-based feature selection algorithm, named GameRank. The method is implemented into an R package composed of additional functions to build calibrated and discriminative predictive models. AVAILABILITY AND IMPLEMENTATION: GameRank is available at https://github.com/Genentech/GameRank and released under the MIT License.
Carsten Henneges, Joseph N. Paulson
Bioinform.2
2021 MicrobiomeExplorer: an R package for the analysis and visualization of microbial communities
abstract
SUMMARY: We developed the MicrobiomeExplorer R package to facilitate the analysis and visualization of microbial communities. The MicrobiomeExplorer R package allows a user to perform typical microbiome analytic workflows and visualize their results, either through the command line or an interactive Shiny application included with the package. In addition to applying common analytical workflows, the application enables automated analysis report generation. AVAILABILITY AND IMPLEMENTATION: Available at https://github.com/zoecastillo/microbiomeExplorer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Janina Reeder, Mo Huang, Joshua S. Kaminker, Joseph N. Paulson
Bioinform.4
2021 Multivariable association discovery in population-scale meta-omics studies
abstract
It is challenging to associate features such as human health outcomes, diet, environmental conditions, or other metadata to microbial community measurements, due in part to their quantitative properties. Microbiome multi-omics are typically noisy, sparse (zero-inflated), high-dimensional, extremely non-normal, and often in the form of count or compositional measurements. Here we introduce an optimized combination of novel and established methodology to assess multivariable association of microbial community features with complex metadata in population-scale observational studies. Our approach, MaAsLin 2 (Microbiome Multivariable Associations with Linear Models), uses generalized linear and mixed models to accommodate a wide variety of modern epidemiological studies, including cross-sectional and longitudinal designs, as well as a variety of data types (e.g., counts and relative abundances) with or without covariates and repeated measurements. To construct this method, we conducted a large-scale evaluation of a broad range of scenarios under which straightforward identification of meta-omics associations can be challenging. These simulation studies reveal that MaAsLin 2's linear model preserves statistical power in the presence of repeated measures and multiple covariates, while accounting for the nuances of meta-omics features and controlling false discovery. We also applied MaAsLin 2 to a microbial multi-omics dataset from the Integrative Human Microbiome (HMP2) project which, in addition to reproducing established results, revealed a unique, integrated landscape of inflammatory bowel diseases (IBD) across multiple time points and omics profiles.
Himel Mallick, Gholamali Rahnavard, Lauren J. McIver, Yancong Zhang, Long H. Nguyen, Timothy L. Tickle, George Weingart, Boyu Ren, Emma H. Schwager, Suvo Chatterjee, Kelsey N. Thompson, Jeremy E. Wilkinson, Ayshwarya Subramanian, Yiren Lu 0003, Levi Waldron, Joseph N. Paulson, Eric A. Franzosa, Héctor Corrada Bravo, Curtis Huttenhower
PLoS Comput. Biol.17
2019 metagenomeFeatures: an R package for working with 16S rRNA reference databases and marker-gene survey feature data
abstract
SUMMARY: We developed the metagenomeFeatures R Bioconductor package along with annotation packages for three 16S rRNA databases (Greengenes, RDP and SILVA) to facilitate working with 16S rRNA databases and marker-gene survey feature data. The metagenomeFeatures package defines two classes, MgDb for working with 16S rRNA sequence databases, and mgFeatures for marker-gene survey feature data. The associated annotation packages provide a consistent interface to the different databases facilitating database comparison and exploration. The mgFeatures-class represents a crucial step in the development of a common data structure for working with 16S marker-gene survey data in R. AVAILABILITY AND IMPLEMENTATION: https://bioconductor.org/packages/release/bioc/html/metagenomeFeatures.html. SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.
Nathan D. Olson, Nidhi Shah, Jayaram Kancherla, Justin Wagner, Joseph N. Paulson, Héctor Corrada Bravo
Bioinform.5
2017 Estimating gene regulatory networks with pandaR
abstract
Abstract PANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that begins with a model of transcription factor–target gene interactions and uses message passing to update the network model given available transcriptomic and protein–protein interaction data. PANDA is used to estimate networks for each experimental group and the network models are then compared between groups to explore transcriptional processes that distinguish the groups. We present pandaR (bioconductor.org/packages/pandaR), a Bioconductor package that implements PANDA and provides a framework for exploratory data analysis on gene regulatory networks. Availability and Implementation: PandaR is provided as a Bioconductor R Package and is available at bioconductor.org/packages/pandaR.
Daniel Schlauch, Joseph N. Paulson, Albert Young, Kimberly Glass, John Quackenbush
Bioinform.2
2017 Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse data
abstract
BACKGROUND: Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. RESULTS: We find that analysis of large RNA-Seq data sets requires both careful quality control and the need to account for sparsity due to the heterogeneity intrinsic in multi-group studies. We developed Yet Another RNA Normalization software pipeline (YARN), that includes quality control and preprocessing, gene filtering, and normalization steps designed to facilitate downstream analysis of large, heterogeneous RNA-Seq data sets and we demonstrate its use with data from the Genotype-Tissue Expression (GTEx) project. CONCLUSIONS: An R package instantiating YARN is available at http://bioconductor.org/packages/yarn .
Joseph N. Paulson, Cho-Yi Chen, Camila Miranda Lopes-Ramos, Marieke L. Kuijjer, John Platig, Abhijeet R. Sonawane, Maud Fagny, Kimberly Glass, John Quackenbush
BMC Bioinform.1
2016 Privacy-preserving microbiome analysis using secure computation
abstract
MOTIVATION: Developing targeted therapeutics and identifying biomarkers relies on large amounts of research participant data. Beyond human DNA, scientists now investigate the DNA of micro-organisms inhabiting the human body. Recent work shows that an individual's collection of microbial DNA consistently identifies that person and could be used to link a real-world identity to a sensitive attribute in a research dataset. Unfortunately, the current suite of DNA-specific privacy-preserving analysis tools does not meet the requirements for microbiome sequencing studies. RESULTS: To address privacy concerns around microbiome sequencing, we implement metagenomic analyses using secure computation. Our implementation allows comparative analysis over combined data without revealing the feature counts for any individual sample. We focus on three analyses and perform an evaluation on datasets currently used by the microbiome research community. We use our implementation to simulate sharing data between four policy-domains. Additionally, we describe an application of our implementation for patients to combine data that allows drug developers to query against and compensate patients for the analysis. AVAILABILITY AND IMPLEMENTATION: The software is freely available for download at: http://cbcb.umd.edu/∼hcorrada/projects/secureseq.html SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Justin Wagner, Joseph N. Paulson, Bobby Bhattacharjee, Héctor Corrada Bravo
Bioinform.2