Michael C. Wu

dblp:52/7064 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
10since 2021 · last 2024
0000-0002-3357-6570ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Statistical analysis of multiple regions-of-interest in multiplexed spatial proteomics data
abstract
Multiplexed spatial proteomics reveals the spatial organization of cells in tumors, which is associated with important clinical outcomes such as survival and treatment response. This spatial organization is often summarized using spatial summary statistics, including Ripley's K and Besag's L. However, if multiple regions of the same tumor are imaged, it is unclear how to synthesize the relationship with a single patient-level endpoint. We evaluate extant approaches for accommodating multiple images within the context of associating summary statistics with outcomes. First, we consider averaging-based approaches wherein multiple summaries for a single sample are combined in a weighted mean. We then propose a novel class of ensemble testing approaches in which we simulate random weights used to aggregate summaries, test for an association with outcomes, and combine the $P$-values. We systematically evaluate the performance of these approaches via simulation and application to data from non-small cell lung cancer, colorectal cancer, and triple negative breast cancer. We find that the optimal strategy varies, but a simple weighted average of the summary statistics based on the number of cells in each image often offers the highest power and controls type I error effectively. When the size of the imaged regions varies, incorporating this variation into the weighted aggregation may yield additional power in cases where the varying size is informative. Ensemble testing (but not resampling) offered high power and type I error control across conditions in our simulated data sets.
Sarah Samorodnitsky, Michael C. Wu
Briefings Bioinform.2
2024 A multi-bin rarefying method for evaluating alpha diversities in TCR sequencing data
abstract
MOTIVATION: T cell receptors (TCRs) constitute a major component of our adaptive immune system, governing the recognition and response to internal and external antigens. Studying the TCR diversity via sequencing technology is critical for a deeper understanding of immune dynamics. However, library sizes differ substantially across samples, hindering the accurate estimation/comparisons of alpha diversities. To address this, researchers frequently use an overall rarefying approach in which all samples are sub-sampled to an even depth. Despite its pervasive application, its efficacy has never been rigorously assessed. RESULTS: In this paper, we develop an innovative "multi-bin" rarefying approach that partitions samples into multiple bins according to their library sizes, conducts rarefying within each bin for alpha diversity calculations, and performs meta-analysis across bins. Extensive simulations using real-world data highlight the inadequacy of the overall rarefying approach in controlling the confounding effect of library size. Our method proves robust in addressing library size confounding, outperforming competing normalization strategies by achieving better-controlled type-I error rates and enhanced statistical power in association tests. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/mli171/MultibinAlpha. The datasets are freely available at https://doi.org/10.21417/B7001Z and https://doi.org/10.21417/AR2019NC.
Xing Hua, Michael C. Wu, Ni Zhao
Bioinform.4
2024 A Spatial Omnibus Test (SPOT) for Spatial Proteomic Data
abstract
MOTIVATION: Spatial proteomics can reveal the spatial organization of immune cells in the tumor immune microenvironment. Relating measures of spatial clustering, such as Ripley's K or Besag's L, to patient outcomes may offer important clinical insights. However, these measures require pre-specifying a radius in which to quantify clustering, yet no consensus exists on the optimal radius which may be context-specific. RESULTS: We propose a SPatial Omnibus Test (SPOT) which conducts this analysis across a range of candidate radii. At each radius, SPOT evaluates the association between the spatial summary and outcome, adjusting for confounders. SPOT then aggregates results across radii using the Cauchy combination test, yielding an omnibus P-value characterizing the overall degree of association. Using simulations, we verify that the type I error rate is controlled and show SPOT can be more powerful than alternatives. We also apply SPOT to ovarian and lung cancer studies. AVAILABILITY AND IMPLEMENTATION: An R package and tutorial are provided at https://github.com/sarahsamorodnitsky/SPOT.
Sarah Samorodnitsky, Katie Campbell, Antoni Ribas, Michael C. Wu
Bioinform.4
2023 Accommodating multiple potential normalizations in microbiome associations studies
abstract
BACKGROUND: Microbial communities are known to be closely related to many diseases, such as obesity and HIV, and it is of interest to identify differentially abundant microbial species between two or more environments. Since the abundances or counts of microbial species usually have different scales and suffer from zero-inflation or over-dispersion, normalization is a critical step before conducting differential abundance analysis. Several normalization approaches have been proposed, but it is difficult to optimize the characterization of the true relationship between taxa and interesting outcomes. RESULTS: To avoid the challenge of picking an optimal normalization and accommodate the advantages of several normalization strategies, we propose an omnibus approach. Our approach is based on a Cauchy combination test, which is flexible and powerful by aggregating individual p values. We also consider a truncated test statistic to prevent substantial power loss. We experiment with a basic linear regression model as well as recently proposed powerful association tests for microbiome data and compare the performance of the omnibus approach with individual normalization approaches. Experimental results show that, regardless of simulation settings, the new approach exhibits power that is close to the best normalization strategy, while controling the type I error well. CONCLUSIONS: The proposed omnibus test releases researchers from choosing among various normalization methods and it is an aggregated method that provides the powerful result to the underlying optimal normalization, which requires tedious trial and error. While the power may not exceed the best normalization, it is always much better than using a poor choice of normalization.
Hoseung Song, Wodan Ling, Ni Zhao, Anna M. Plantinga, Courtney A. Broedlow, Nichole R. Klatt, Tiffany Hensley-McBain, Michael C. Wu
BMC Bioinform.8
2022 Testing microbiome association using integrated quantile regression models
abstract
MOTIVATION: Most existing microbiome association analyses focus on the association between microbiome and conditional mean of health or disease-related outcomes, and within this vein, vast computational tools and methods have been devised for standard binary or continuous outcomes. However, these methods tend to be limited either when the underlying microbiome-outcome association occurs somewhere other than the mean level, or when distribution of the outcome variable is irregular (e.g. zero-inflated or mixtures) such that conditional outcome mean is less meaningful. We address this gap by investigating association analysis between microbiome compositions and conditional outcome quantiles. RESULTS: We introduce a new association analysis tool named MiRKAT-IQ within the Microbiome Regression-based Kernel Association Test framework using Integrated Quantile regression models to examine the association between microbiome and the distribution of outcome. For an individual quantile, we utilize the existing kernel machine regression framework to examine the association between that conditional outcome quantile and a group of microbial features (e.g. microbiome community compositions). Then, the goal of examining microbiome association with the whole outcome distribution is achieved by integrating all outcome conditional quantiles over a process, and thus our new MiRKAT-IQ test is robust to both the location of association signals (e.g. mean, variance, median) and the heterogeneous distribution of the outcome. Extensive numerical simulation studies have been conducted to show the validity of the new MiRKAT-IQ test. We demonstrate the potential usefulness of MiRKAT-IQ with applications to actual biological data collected from a previous microbiome study. AVAILABILITY AND IMPLEMENTATION: R codes to implement the proposed methodology is provided in the MiRKAT package, which is available on CRAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianying Wang, Wodan Ling, Anna M. Plantinga, Michael C. Wu, Xiang Zhan
Bioinform.4
2022 TCR-L: an analysis tool for evaluating the association between the T-cell receptor repertoire and clinical phenotypes
abstract
BACKGROUND: T cell receptors (TCRs) play critical roles in adaptive immune responses, and recent advances in genome technology have made it possible to examine the T cell receptor (TCR) repertoire at the individual sequence level. The analysis of the TCR repertoire with respect to clinical phenotypes can yield novel insights into the etiology and progression of immune-mediated diseases. However, methods for association analysis of the TCR repertoire have not been well developed. METHODS: We introduce an analysis tool, TCR-L, for evaluating the association between the TCR repertoire and disease outcomes. Our approach is developed under a mixed effect modeling, where the fixed effect represents features that can be explicitly extracted from TCR sequences while the random effect represents features that are hidden in TCR sequences and are difficult to be extracted. Statistical tests are developed to examine the two types of effects independently, and then the p values are combined. RESULTS: Simulation studies demonstrate that (1) the proposed approach can control the type I error well; and (2) the power of the proposed approach is greater than approaches that consider fixed effect only or random effect only. The analysis of real data from a skin cutaneous melanoma study identifies an association between the TCR repertoire and the short/long-term survival of patients. CONCLUSION: The TCR-L can accommodate features that can be extracted as well as features that are hidden in TCR sequences. TCR-L provides a powerful approach for identifying association between TCR repertoire and disease outcomes.
Juna Goo, Yang Liu 0312, Michael C. Wu, Li Hsu, Qianchuan He
BMC Bioinform.5
2021 Deep ensemble learning over the microbial phylogenetic tree (DeepEn-Phy)
abstract
Successful prediction of clinical outcomes facilitates tailored diagnosis and treatment. The microbiome has been shown to be an important biomarker to predict host clinical outcomes. Further, the incorporation of microbial phylogeny, the evolutionary relationship among microbes, has been demonstrated to improve prediction accuracy. We propose a phylogeny-driven deep neural network (PhyNN) and develop an ensemble method, DeepEn-Phy, for host clinical outcome prediction. The method is designed to optimally extract features from phylogeny, thereby take full advantage of the information in phylogeny while harnessing the core principles of phylogeny (in contrast to taxonomy). We apply DeepEn-Phy to a real large microbiome data set to predict both categorical and continuous clinical outcomes. DeepEn-Phy demonstrates superior prediction performance to existing machine learning and deep learning approaches. Overall, DeepEn-Phy provides a new strategy for designing deep neural network architectures within the context of phylogeny-constrained microbiome data.
Wodan Ling, Youran Qi, Xing Hua, Michael C. Wu
BIBM4
2021 A Kernel-based Test of Independence for Cluster-correlated Data
abstract
The Hilbert-Schmidt Independence Criterion (HSIC) is a powerful kernel-based statistic for assessing the generalized dependence between two multivariate variables. However, independence testing based on the HSIC is not directly possible for cluster-correlated data. Such a correlation pattern among the observations arises in many practical situations, e.g., family-based and longitudinal data, and requires proper accommodation. Therefore, we propose a novel HSIC-based independence test to evaluate the dependence between two multivariate variables based on cluster-correlated data. Using the previously proposed empirical HSIC as our test statistic, we derive its asymptotic distribution under the null hypothesis of independence between the two variables but in the presence of sample correlation. Based on both simulation studies and real data analysis, we show that, with clustered data, our approach effectively controls type I error and has a higher statistical power than competing methods.
Hongjiao Liu, Anna M. Plantinga, Yunhua Xiang, Michael C. Wu
NeurIPS4
2021 A method for subtype analysis with somatic mutations
abstract
MOTIVATION: Cancer is a highly heterogeneous disease, and virtually all types of cancer have subtypes. Understanding the association between cancer subtypes and genetic variations is fundamental to the development of targeted therapies for patients. Somatic mutation plays important roles in tumor development and has emerged as a new type of genetic variations for studying the association with cancer subtypes. However, the low prevalence of individual mutations poses a tremendous challenge to the related statistical analysis. RESULTS: In this article, we propose an approach, subtype analysis with somatic mutations (SASOM), for the association analysis of cancer subtypes with somatic mutations. Our approach tests the association between a set of somatic mutations (from a genetic pathway) and subtypes, while incorporating functional information of the mutations into the analysis. We further propose a robust p-value combination procedure, DAPC, to synthesize statistical significance from different sources. Simulation studies show that the proposed approach has correct type I error and tends to be more powerful than possible alternative methods. In a real data application, we examine the somatic mutations from a cutaneous melanoma dataset, and identify a genetic pathway that is associated with immune-related subtypes. AVAILABILITY AND IMPLEMENTATION: The SASOM R package is available at https://github.com/rksyouyou/SASOM-pkg. R scripts and data are available at https://github.com/rksyouyou/SASOM-analysis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yang Liu 0312, Michael C. Wu, Li Hsu, Qianchuan He
Bioinform.3
2021 MiRKAT: kernel machine regression-based global association tests for the microbiome
abstract
SUMMARY: Distance-based tests of microbiome beta diversity are an integral part of many microbiome analyses. MiRKAT enables distance-based association testing with a wide variety of outcome types, including continuous, binary, censored time-to-event, multivariate, correlated and high-dimensional outcomes. Omnibus tests allow simultaneous consideration of multiple distance and dissimilarity measures, providing higher power across a range of simulation scenarios. Two measures of effect size, a modified R-squared coefficient and a kernel RV coefficient, are incorporated to allow comparison of effect sizes across multiple kernels. AVAILABILITY AND IMPLEMENTATION: MiRKAT is available on CRAN as an R package. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nehemiah Wilson, Ni Zhao, Xiang Zhan, Hyunwook Koh, Weijia Fu, Hongzhe Li, Michael C. Wu, Anna M. Plantinga
Bioinform.8
2019 pldist: ecological dissimilarities for paired and longitudinal microbiome association analysis
abstract
MOTIVATION: The human microbiome is notoriously variable across individuals, with a wide range of 'healthy' microbiomes. Paired and longitudinal studies of the microbiome have become increasingly popular as a way to reduce unmeasured confounding and to increase statistical power by reducing large inter-subject variability. Statistical methods for analyzing such datasets are scarce. RESULTS: We introduce a paired UniFrac dissimilarity that summarizes within-individual (or within-pair) shifts in microbiome composition and then compares these compositional shifts across individuals (or pairs). This dissimilarity depends on a novel transformation of relative abundances, which we then extend to more than two time points and incorporate into several phylogenetic and non-phylogenetic dissimilarities. The data transformation and resulting dissimilarities may be used in a wide variety of downstream analyses, including ordination analysis and distance-based hypothesis testing. Simulations demonstrate that tests based on these dissimilarities retain appropriate type 1 error and high power. We apply the method in two real datasets. AVAILABILITY AND IMPLEMENTATION: The R package pldist is available on GitHub at https://github.com/aplantin/pldist. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anna M. Plantinga, Robert R. Jenq, Michael C. Wu
Bioinform.4
2019 Identifying individual risk rare variants using protein structure guided local tests (POINT)
abstract
Rare variants are of increasing interest to genetic association studies because of their etiological contributions to human complex diseases. Due to the rarity of the mutant events, rare variants are routinely analyzed on an aggregate level. While aggregation analyses improve the detection of global-level signal, they are not able to pinpoint causal variants within a variant set. To perform inference on a localized level, additional information, e.g., biological annotation, is often needed to boost the information content of a rare variant. Following the observation that important variants are likely to cluster together on functional domains, we propose a protein structure guided local test (POINT) to provide variant-specific association information using structure-guided aggregation of signal. Constructed under a kernel machine framework, POINT performs local association testing by borrowing information from neighboring variants in the 3-dimensional protein space in a data-adaptive fashion. Besides merely providing a list of promising variants, POINT assigns each variant a p-value to permit variant ranking and prioritization. We assess the selection performance of POINT using simulations and illustrate how it can be used to prioritize individual rare variants in PCSK9, ANGPTL4 and CETP in the Action to Control Cardiovascular Risk in Diabetes (ACCORD) clinical trial data.
Rachel Marceau West, Wenbin Lu, Daniel M. Rotroff, Mélaine A. Kuenemann, Sheng-Mao Chang, Michael C. Wu, Michael J. Wagner, John B. Buse, Alison A. Motsinger-Reif, Denis Fourches, Jung-Ying Tzeng
PLoS Comput. Biol.6
2016 A novel copy number variants kernel association test with application to autism spectrum disorders studies
abstract
MOTIVATION: Copy number variants (CNVs) have been implicated in a variety of neurodevelopmental disorders, including autism spectrum disorders, intellectual disability and schizophrenia. Recent advances in high-throughput genomic technologies have enabled rapid discovery of many genetic variants including CNVs. As a result, there is increasing interest in studying the role of CNVs in the etiology of many complex diseases. Despite the availability of an unprecedented wealth of CNV data, methods for testing association between CNVs and disease-related traits are still under-developed due to the low prevalence and complicated multi-scale features of CNVs. RESULTS: We propose a novel CNV kernel association test (CKAT) in this paper. To address the low prevalence, CNVs are first grouped into CNV regions (CNVR). Then, taking into account the multi-scale features of CNVs, we first design a single-CNV kernel which summarizes the similarity between two CNVs, and next aggregate the single-CNV kernel to a CNVR kernel which summarizes the similarity between two CNVRs. Finally, association between CNVR and disease-related traits is assessed by comparing the kernel-based similarity with the similarity in the trait using a score test for variance components in a random effect model. We illustrate the proposed CKAT using simulations and show that CKAT is more powerful than existing methods, while always being able to control the type I error. We also apply CKAT to a real dataset examining the association between CNV and autism spectrum disorders, which demonstrates the potential usefulness of the proposed method. AVAILABILITY AND IMPLEMENTATION: A R package to implement the proposed CKAT method is available at http://works.bepress.com/debashis_ghosh/ CONTACTS: [email protected] or [email protected] information: Supplementary data are available at Bioinformatics online.
Xiang Zhan, Santhosh Girirajan, Ni Zhao, Michael C. Wu, Debashis Ghosh
Bioinform.4
2009 Sparse linear discriminant analysis for simultaneous testing for the significance of a gene set/pathway and gene selection
abstract
MOTIVATION: Pathway and gene set-based approaches for the analysis of gene expression profiling experiments have become increasingly popular for addressing problems associated with individual gene analysis. Since most genes are not differently expressed, existing gene set tests, which consider all the genes within a gene set, are subject to considerable noise and power loss, a concern exacerbated in studies in which the degree of differential expression is moderate for truly differentially expressed genes. For a significantly differentially expressed pathway, it is also of substantial interest to select important genes that drive the differential expression of the pathway. METHODS: We develop a unified framework to jointly test the significance of a pathway and to select a subset of genes that drive the significant pathway effect. To achieve dimension reduction and gene selection, we decompose each gene pathway into a single score by using a regularized form of linear discriminant analysis, called sparse linear discriminant analysis (sLDA). Testing for the significance of the pathway effect proceeds via permutation of the sLDA score. The sLDA-based test is compared with competing approaches with simulations and two applications: a study on the effect of metal fume exposure on immune response and a study of gene expression profiles among Type II Diabetes patients. RESULTS: Our results show that sLDA-based testing provides a powerful approach to test for the significance of a differentially expressed pathway and gene selection. AVAILABILITY: An implementation of the proposed sLDA-based pathway test in the R statistical computing environment is available at http://www.hsph.harvard.edu/~mwu/software/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Michael C. Wu, Lingsong Zhang, Zhaoxi Wang, David Christiani, Xihong Lin
Bioinform.1