VLDB 2026 Research / reviewers in the wild / expert
Christine B. Peterson
dblp:184/4912
· DBLP profile ↗
9ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-3316-0468ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | B-MASTER: scalable Bayesian multivariate regression for master predictor discovery in colorectal cancer microbiome-metabolite profilesabstractMOTIVATION: The gut microbiome shapes cancer therapy response through its influence on host metabolism. While prior studies examine pairwise associations between individual genera and metabolites, there is limited methodology for identifying microbial genera that systematically regulate the overall metabolome. Scalable statistical tools are needed to uncover such system-level "master predictors" in high-dimensional microbiome-metabolome data. RESULTS: We introduce B-MASTER, a scalable Bayesian multivariate regression framework combining ℓ1 sparsity and ℓ2 group shrinkage to identify essential cross-metabolite regulators. A Gibbs sampler enables near-linear computational scaling, supporting models with millions of parameters. The method is supported by theoretical guarantees, including posterior contraction and selection consistency. Analysis of colorectal cancer microbiome-metabolome data reveals key microbial genera that govern global and cancer-associated metabolite patterns, highlighting system-level regulatory structure. AVAILABILITY AND IMPLEMENTATION: The B-MASTER code, including demonstration scripts, is available at https://github.com/priyamdas2/B-MASTER. An archived snapshot of the code corresponding to this manuscript is available on Zenodo with DOI: 10.5281/zenodo.20484958. Priyam Das, Tanujit Dey, Christine B. Peterson, Sounak Chakraborty |
Bioinform. | 3 |
| 2025 | CAT: a conditional association test for microbiome data using a permutation approachabstractIn microbiome analysis, researchers often seek to identify taxonomic features associated with an outcome of interest. However, microbiome features are intercorrelated and linked by phylogenetic relationships, making it challenging to assess the association between an individual feature and an outcome. This paper proposes a novel conditional association test, CAT, that can account for other features and phylogenetic relatedness when testing the association between a feature and an outcome. CAT adopts a permutation approach, measuring the importance of a feature in predicting the outcome by permuting operational taxonomic unit/amplicon sequence variant counts belonging to that feature from the data and quantifying how much the association with the outcome is weakened through the change in the coefficient of determination $R^{2}$. Compared with marginal association tests, it focuses on the added value of a feature in explaining outcome variation that is not captured by other features. By leveraging global tests including PERMANOVA and MiRKAT-based methods, CAT allows association testing for continuous, binary, categorical, count, survival, and correlated outcomes. We demonstrate through simulation studies that CAT can provide a direct quantification of feature importance that is distinct from that of marginal association tests, and illustrate CAT with applications to two real-world studies on the microbiome in melanoma patients: one examining the role of the microbiome in shaping immunotherapy response, and one investigating the association between the microbiome and survival outcomes. Our results illustrate the potential of CAT to inform the design of microbiome interventions aimed at improving clinical outcomes. Yushu Shi, Kim-Anh Do, Robert R. Jenq, Christine B. Peterson |
Briefings Bioinform. | 5 |
| 2024 | TARO: tree-aggregated factor regression for microbiome data integrationabstractMOTIVATION: Although the human microbiome plays a key role in health and disease, the biological mechanisms underlying the interaction between the microbiome and its host are incompletely understood. Integration with other molecular profiling data offers an opportunity to characterize the role of the microbiome and elucidate therapeutic targets. However, this remains challenging to the high dimensionality, compositionality, and rare features found in microbiome profiling data. These challenges necessitate the use of methods that can achieve structured sparsity in learning cross-platform association patterns. RESULTS: We propose Tree-Aggregated factor RegressiOn (TARO) for the integration of microbiome and metabolomic data. We leverage information on the taxonomic tree structure to flexibly aggregate rare features. We demonstrate through simulation studies that TARO accurately recovers a low-rank coefficient matrix and identifies relevant features. We applied TARO to microbiome and metabolomic profiles gathered from subjects being screened for colorectal cancer to understand how gut microrganisms shape intestinal metabolite abundances. AVAILABILITY AND IMPLEMENTATION: The R package TARO implementing the proposed methods is available online at https://github.com/amishra-stats/taro-package. Aditya K. Mishra, Iqbal Mahmud, Philip L. Lorenzi, Robert R. Jenq, Jennifer A. Wargo, Nadim J. Ajami, Christine B. Peterson |
Bioinform. | 7 |
| 2022 | Estimating the optimal linear combination of predictors using spherically constrained optimizationabstractIn the context of a binary classification problem, the optimal linear combination of continuous predictors can be estimated by maximizing an empirical estimate of the area under the receiver operating characteristic (ROC) curve (AUC). For multi-category responses, the optimal predictor combination can similarly be obtained by maximization of the empirical hypervolume under the manifold (HUM). This problem is particularly relevant to medical research, where it may be of interest to diagnose a disease with various subtypes or predict a multi-category outcome. Since the empirical HUM is discontinuous, non-differentiable, and possibly multi-modal, solving this maximization problem requires a global optimization technique. Estimation of the optimal coefficient vector using existing global optimization techniques is computationally expensive, becoming prohibitive as the number of predictors and the number of outcome categories increases. We propose an efficient derivative-free black-box optimization technique based on pattern search to solve this problem. Through extensive simulation studies, we demonstrate that the proposed method achieves better performance compared to existing methods including the step-down algorithm. Finally, we illustrate the proposed method to predict swallowing difficulty after radiation therapy for oropharyngeal cancer based on radiation dose to various structures in the head and neck. Priyam Das, Debsurya De, Raju Maiti, Mona Kamal, Katherine A. Hutcheson, Clifton D. Fuller, Bibhas Chakraborty, Christine B. Peterson |
BMC Bioinform. | 8 |
| 2021 | ProgPerm: Progressive permutation for a dynamic representation of the robustness of microbiome discoveriesabstractBACKGROUND: Identification of features is a critical task in microbiome studies that is complicated by the fact that microbial data are high dimensional and heterogeneous. Masked by the complexity of the data, the problem of separating signals (differential features between groups) from noise (features that are not differential between groups) becomes challenging and troublesome. For instance, when performing differential abundance tests, multiple testing adjustments tend to be overconservative, as the probability of a type I error (false positive) increases dramatically with the large numbers of hypotheses. Moreover, the grouping effect of interest can be obscured by heterogeneity. These factors can incorrectly lead to the conclusion that there are no differences in the microbiome compositions. RESULTS: We translate and represent the problem of identifying differential features, which are differential in two-group comparisons (e.g., treatment versus control), as a dynamic layout of separating the signal from its random background. More specifically, we progressively permute the grouping factor labels of the microbiome samples and perform multiple differential abundance tests in each scenario. We then compare the signal strength of the most differential features from the original data with their performance in permutations, and will observe a visually apparent decreasing trend if these features are true positives identified from the data. Simulations and applications on real data show that the proposed method creates a U-curve when plotting the number of significant features versus the proportion of mixing. The shape of the U-Curve can convey the strength of the overall association between the microbiome and the grouping factor. We also define a fragility index to measure the robustness of the discoveries. Finally, we recommend the identified features by comparing p-values in the observed data with p-values in the fully mixed data. CONCLUSIONS: We have developed this into a user-friendly and efficient R-shiny tool with visualizations. By default, we use the Wilcoxon rank sum test to compute the p-values, since it is a robust nonparametric test. Our proposed method can also utilize p-values obtained from other testing methods, such as DESeq. This demonstrates the potential of the progressive permutation method to be extended to new settings. Yushu Shi, Kim-Anh Do, Christine B. Peterson, Robert R. Jenq |
BMC Bioinform. | 4 |
| 2020 | NExUS: Bayesian simultaneous network estimation across unequal sample sizesabstractMOTIVATION: Network-based analyses of high-throughput genomics data provide a holistic, systems-level understanding of various biological mechanisms for a common population. However, when estimating multiple networks across heterogeneous sub-populations, varying sample sizes pose a challenge in the estimation and inference, as network differences may be driven by differences in power. We are particularly interested in addressing this challenge in the context of proteomic networks for related cancers, as the number of subjects available for rare cancer (sub-)types is often limited. RESULTS: We develop NExUS (Network Estimation across Unequal Sample sizes), a Bayesian method that enables joint learning of multiple networks while avoiding artefactual relationship between sample size and network sparsity. We demonstrate through simulations that NExUS outperforms existing network estimation methods in this context, and apply it to learn network similarity and shared pathway activity for groups of cancers with related origins represented in The Cancer Genome Atlas (TCGA) proteomic data. AVAILABILITY AND IMPLEMENTATION: The NExUS source code is freely available for download at https://github.com/priyamdas2/NExUS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Priyam Das, Christine B. Peterson, Kim-Anh Do, Rehan Akbani, Veerabhadran Baladandayuthapani |
Bioinform. | 2 |
| 2020 | aPCoA: covariate adjusted principal coordinates analysisabstractSUMMARY: In fields, such as ecology, microbiology and genomics, non-Euclidean distances are widely applied to describe pairwise dissimilarity between samples. Given these pairwise distances, principal coordinates analysis is commonly used to construct a visualization of the data. However, confounding covariates can make patterns related to the scientific question of interest difficult to observe. We provide adjusted principal coordinates analysis as an easy-to-use tool, available as both an R package and a Shiny app, to improve data visualization in this context, enabling enhanced presentation of the effects of interest. AVAILABILITY AND IMPLEMENTATION: The R package 'aPCoA' and Shiny app can be accessed at https://cran.r-project.org/web/packages/aPCoA/index.html and https://biostatistics.mdanderson.org/shinyapps/aPCoA/. Yushu Shi, Kim-Anh Do, Christine B. Peterson, Robert R. Jenq |
Bioinform. | 4 |
| 2020 | Compositional zero-inflated network estimation for microbiome dataabstractBACKGROUND: The estimation of microbial networks can provide important insight into the ecological relationships among the organisms that comprise the microbiome. However, there are a number of critical statistical challenges in the inference of such networks from high-throughput data. Since the abundances in each sample are constrained to have a fixed sum and there is incomplete overlap in microbial populations across subjects, the data are both compositional and zero-inflated. RESULTS: We propose the COmpositional Zero-Inflated Network Estimation (COZINE) method for inference of microbial networks which addresses these critical aspects of the data while maintaining computational scalability. COZINE relies on the multivariate Hurdle model to infer a sparse set of conditional dependencies which reflect not only relationships among the continuous values, but also among binary indicators of presence or absence and between the binary and continuous representations of the data. Our simulation results show that the proposed method is better able to capture various types of microbial relationships than existing approaches. We demonstrate the utility of the method with an application to understanding the oral microbiome network in a cohort of leukemic patients. CONCLUSIONS: Our proposed method addresses important challenges in microbiome network estimation, and can be effectively applied to discover various types of dependence relationships in microbial communities. The procedure we have developed, which we refer to as COZINE, is available online at https://github.com/MinJinHa/COZINE . Min Jin Ha, Junghi Kim, Jessica Galloway-Pena, Kim-Anh Do, Christine B. Peterson |
BMC Bioinform. | 5 |
| 2016 | TreeQTL: hierarchical error control for eQTL findingsabstractUNLABELLED: : Commonly used multiplicity adjustments fail to control the error rate for reported findings in many expression quantitative trait loci (eQTL) studies. TreeQTL implements a hierarchical multiple testing procedure which allows control of appropriate error rates defined relative to a grouping of the eQTL hypotheses. AVAILABILITY AND IMPLEMENTATION: The R package TreeQTL is available for download at http://bioinformatics.org/treeqtl CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Christine B. Peterson, M. Bogomolov, Yoav Benjamini, Chiara Sabatti |
Bioinform. | 1 |