EDBT 2026 Demo / reviewers in the wild / expert
Glen A. Satten
dblp:66/6945
· DBLP profile ↗
10ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-7275-5371ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTH : a nonparametric test for homogeneity of multivariate dispersionsabstractAbstract Motivation Testing for differences in within-group dispersion is a fundamental problem in multivariate data analysis, with direct implications for interpreting group structure and validating statistical assumptions of other analysis such as ANOVA. Existing methods typically construct test statistics either based on the distance of each observation from the group center or on the mean of pairwise dissimilarities among observations within a group. Both approaches can fail when the mean within-group distance is similar across groups but the distributions of the within-group distances differ. This issue is particularly relevant in high-dimensional microbiome data, where outliers and overdispersion can distort the performance of mean-dissimilarity-based tests. Results We introduce the non-parametric Distance-based Test for Homogeneity (DTH), which measures dispersion of a group by computing within-group dissimilarity. Difference in dispersion across groups is tested by comparing the distributions of the within-group dissimilarity across different groups. A combination of Kolmogorov-Smirnov and Wasserstein distances are used to construct the difference between the distributions. For more than two groups, pairwise group tests are combined using a permutation-based p-value. Through simulations, we show that our method has higher power than existing tests for homogeneity in certain situations and comparable power in others. For continuous covariates, we offer an heuristic extension of DTH that showed good performance in simulations. Availability and implementation The DTH package, along with the code for reproducing all simulations, analyses, and an accompanying vignette, is available at https://github.com/asmita112358/DTH. Asmita Roy, Jiuyao Lu, Glen A. Satten, Ni Zhao |
Bioinform. | 3 |
| 2023 | Compositional analysis of microbiome data using the linear decomposition model (LDM)abstractSUMMARY: There are compelling reasons to test compositional hypotheses about microbiome data. We present here linear decomposition model-centered log ratio (LDM-clr), an extension of our LDM approach to allow fitting linear models to centered-log-ratio-transformed taxa count data. As LDM-clr is implemented within the existing LDM program, this extension enjoys all the features supported by LDM, including a compositional analysis of differential abundance at both the taxon and community levels, while allowing for a wide range of covariates and study designs for either association or mediation analysis. AVAILABILITY AND IMPLEMENTATION: LDM-clr has been added to the R package LDM, which is available on GitHub at https://github.com/yijuanhu/LDM. Yi-Juan Hu, Glen A. Satten |
Bioinform. | 2 |
| 2022 | A rarefaction-without-resampling extension of PERMANOVA for testing presence-absence associations in the microbiomeabstractMOTIVATION: PERMANOVA is currently the most commonly used method for testing community-level hypotheses about microbiome associations with covariates of interest. PERMANOVA can test for associations that result from changes in which taxa are present or absent by using the Jaccard or unweighted UniFrac distance. However, such presence-absence analyses face a unique challenge: confounding by library size (total sample read count), which occurs when library size is associated with covariates in the analysis. It is known that rarefaction (subsampling to a common library size) controls this bias but at the potential costs of information loss and the introduction of a stochastic component into the analysis. RESULTS: Here, we develop a non-stochastic approach to PERMANOVA presence-absence analyses that aggregates information over all potential rarefaction replicates without actual resampling, when the Jaccard or unweighted UniFrac distance is used. We compare this new approach to three possible ways of aggregating PERMANOVA over multiple rarefactions obtained from resampling: averaging the distance matrix, averaging the (element-wise) squared distance matrix and averaging the F-statistic. Our simulations indicate that our non-stochastic approach is robust to confounding by library size and outperforms each of the stochastic resampling approaches. We also show that, when overdispersion is low, averaging the (element-wise) squared distance outperforms averaging the unsquared distance, currently implemented in the R package vegan. We illustrate our methods using an analysis of data on inflammatory bowel disease in which samples from case participants have systematically smaller library sizes than samples from control participants. AVAILABILITY AND IMPLEMENTATION: We have implemented all the approaches described above, including the function for calculating the analytical average of the squared or unsquared distance matrix, in our R package LDM, which is available on GitHub at https://github.com/yijuanhu/LDM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi-Juan Hu, Glen A. Satten |
Bioinform. | 2 |
| 2022 | Integrative analysis of relative abundance data and presence-absence data of the microbiome using the LDMabstractSUMMARY: We previously developed the LDM for testing hypotheses about the microbiome that performs the test at both the community level and the individual taxon level. The LDM can be applied to relative abundance data and presence-absence data separately, which work well when associated taxa are abundant and rare, respectively. Here, we propose LDM-omni3 that combines LDM analyses at the relative abundance and presence-absence data scales, thereby offering optimal power across scenarios with different association mechanisms. The new LDM-omni3 test is available for the wide range of data types and analyses that are supported by the LDM. AVAILABILITY AND IMPLEMENTATION: The LDM-omni3 test has been added to the R package LDM, which is available on GitHub at https://github.com/yijuanhu/LDM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhengyi Zhu, Glen A. Satten, Yi-Juan Hu |
Bioinform. | 2 |
| 2022 | Testing microbiome associations with survival times at both the community and individual taxon levelsabstractBACKGROUND: Finding microbiome associations with possibly censored survival times is an important problem, especially as specific taxa could serve as biomarkers for disease prognosis or as targets for therapeutic interventions. The two existing methods for survival outcomes, MiRKAT-S and OMiSA, are restricted to testing associations at the community level and do not provide results at the individual taxon level. An ad hoc approach testing each taxon with a survival outcome using the Cox proportional hazard model may not perform well in the microbiome setting with sparse count data and small sample sizes. METHODS: We have previously developed the linear decomposition model (LDM) for testing continuous or discrete outcomes that unifies community-level and taxon-level tests into one framework. Here we extend the LDM to test survival outcomes. We propose to use the Martingale residuals or the deviance residuals obtained from the Cox model as continuous covariates in the LDM. We further construct tests that combine the results of analyzing each set of residuals separately. Finally, we extend PERMANOVA, the most commonly used distance-based method for testing community-level hypotheses, to handle survival outcomes in a similar manner. RESULTS: Using simulated data, we showed that the LDM-based tests preserved the false discovery rate for testing individual taxa and had good sensitivity. The LDM-based community-level tests and PERMANOVA-based tests had comparable or better power than MiRKAT-S and OMiSA. An analysis of data on the association of the gut microbiome and the time to acute graft-versus-host disease revealed several dozen associated taxa that would not have been achievable by any community-level test, as well as improved community-level tests by the LDM and PERMANOVA over those obtained using MiRKAT-S and OMiSA. CONCLUSIONS: Unlike existing methods, our new methods are capable of discovering individual taxa that are associated with survival times, which could be of important use in clinical settings. Yingtian Hu, Glen A. Satten, Yi-Juan Hu |
PLoS Comput. Biol. | 3 |
| 2021 | A rarefaction-based extension of the LDM for testing presence-absence associations in the microbiomeabstractMOTIVATION: Many methods for testing association between the microbiome and covariates of interest (e.g. clinical outcomes, environmental factors) assume that these associations are driven by changes in the relative abundance of taxa. However, these associations may also result from changes in which taxa are present and which are absent. Analyses of such presence-absence associations face a unique challenge: confounding by library size (total sample read count), which occurs when library size is associated with covariates in the analysis. It is known that rarefaction (subsampling to a common library size) controls this bias, but at the potential cost of information loss as well as the introduction of a stochastic component into the analysis. Currently, there is a need for robust and efficient methods for testing presence-absence associations in the presence of such confounding, both at the community level and at the individual-taxon level, that avoid the drawbacks of rarefaction. RESULTS: We have previously developed the linear decomposition model (LDM) that unifies the community-level and taxon-level tests into one framework. Here, we present an extension of the LDM for testing presence-absence associations. The extended LDM is a non-stochastic approach that repeatedly applies the LDM to all rarefied taxa count tables, averages the residual sum-of-squares (RSS) terms over the rarefaction replicates, and then forms an F-statistic based on these average RSS terms. We show that this approach compares favorably to averaging the F-statistic from R rarefaction replicates, which can only be calculated stochastically. The flexible nature of the LDM allows discrete or continuous traits or interactions to be tested while allowing confounding covariates to be adjusted for. Our simulations indicate that our proposed method is robust to any systematic differences in library size and has better power than alternative approaches. We illustrate our method using an analysis of data on inflammatory bowel disease (IBD) in which cases have systematically smaller library sizes than controls. AVAILABILITYAND IMPLEMENTATION: The R package LDM is available on GitHub at https://github.com/yijuanhu/LDM in formats appropriate for Macintosh or Windows. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi-Juan Hu, Andrea Lane, Glen A. Satten |
Bioinform. | 3 |
| 2020 | Testing hypotheses about the microbiome using the linear decomposition model (LDM)abstractMOTIVATION: Methods for analyzing microbiome data generally fall into one of two groups: tests of the global hypothesis of any microbiome effect, which do not provide any information on the contribution of individual operational taxonomic units (OTUs); and tests for individual OTUs, which do not typically provide a global test of microbiome effect. Without a unified approach, the findings of a global test may be hard to resolve with the findings at the individual OTU level. Further, many tests of individual OTU effects do not preserve the false discovery rate (FDR). RESULTS: We introduce the linear decomposition model (LDM), that provides a single analysis path that includes global tests of any effect of the microbiome, tests of the effects of individual OTUs while accounting for multiple testing by controlling the FDR, and a connection to distance-based ordination. The LDM accommodates both continuous and discrete variables (e.g. clinical outcomes, environmental factors) as well as interaction terms to be tested either singly or in combination, allows for adjustment of confounding covariates, and uses permutation-based P-values that can control for sample correlation. The LDM can also be applied to transformed data, and an 'omnibus' test can easily combine results from analyses conducted on different transformation scales. We also provide a new implementation of PERMANOVA based on our approach. For global testing, our simulations indicate the LDM provided correct type I error and can have comparable power to existing distance-based methods. For testing individual OTUs, our simulations indicate the LDM controlled the FDR well. In contrast, DESeq2 often had inflated FDR; MetagenomeSeq generally had the lowest sensitivity. The flexibility of the LDM for a variety of microbiome studies is illustrated by the analysis of data from two microbiome studies. We also show that our implementation of PERMANOVA can outperform existing implementations. AVAILABILITY AND IMPLEMENTATION: The R package LDM is available on GitHub at https://github.com/yijuanhu/LDM in formats appropriate for Macintosh or Windows. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi-Juan Hu, Glen A. Satten |
Bioinform. | 2 |
| 2018 | Robust inference of population structure from next-generation sequencing data with systematic differences in sequencingabstractMotivation: Inferring population structure is important for both population genetics and genetic epidemiology. Principal components analysis (PCA) has been effective in ascertaining population structure with array genotype data but can be difficult to use with sequencing data, especially when low depth leads to uncertainty in called genotypes. Because PCA is sensitive to differences in variability, PCA using sequencing data can result in components that correspond to differences in sequencing quality (read depth and error rate), rather than differences in population structure. We demonstrate that even existing methods for PCA specifically designed for sequencing data can still yield biased conclusions when used with data having sequencing properties that are systematically different across different groups of samples (i.e. sequencing groups). This situation can arise in population genetics when combining sequencing data from different studies, or in genetic epidemiology when using historical controls such as samples from the 1000 Genomes Project. Results: To allow inference on population structure using PCA in these situations, we provide an approach that is based on using sequencing reads directly without calling genotypes. Our approach is to adjust the data from different sequencing groups to have the same read depth and error rate so that PCA does not generate spurious components representing sequencing quality. To accomplish this, we have developed a subsampling procedure to match the depth distributions in different sequencing groups, and a read-flipping procedure to match the error rates. We average over subsamples and read flips to minimize loss of information. We demonstrate the utility of our approach using two datasets from 1000 Genomes, and further evaluate it using simulation studies. Availability and implementation: TASER-PC software is publicly available at http://web1.sph.emory.edu/users/yhu30/software.html. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Peizhou Liao, Glen A. Satten, Yi-Juan Hu |
Bioinform. | 2 |
| 2004 | An empirical bayes adjustment to increase the sensitivity of detecting differentially expressed genes in microarray experimentsabstractMOTIVATION: Detection of differentially expressed genes is one of the major goals of microarray experiments. Pairwise comparison for each gene is not appropriate without controlling the overall (experimentwise) type 1 error rate. Dudoit et al. have advocated use of permutation-based step-down P-value adjustments to correct the observed significance levels for the individual (i.e. for each gene) two sample t-tests. RESULTS: In this paper, we consider an ANOVA formulation of the gene expression levels corresponding to multiple tissue types. We provide resampling-based step-down adjustments to correct the observed significance levels for the individual ANOVA t-tests for each gene and for each pair of tissue type comparisons. More importantly, we introduce a novel empirical Bayes adjustment to the t-test statistics that can be incorporated into the step-down procedure. Using simulated data, we show that the empirical Bayes adjustment improved the sensitivity of detecting differentially expressed genes up to 16%, while maintaining a high level of specificity. This adjustment also reduces the false non-discovery rate to some degree at the cost of a modest increase in the false discovery rate. We illustrate our approach using a human colon cancer dataset consisting of oligonucleotide arrays of normal, adenoma and carcinoma cells. The number of genes with differential expression level declared statistically significant was about 50 when comparing normal to adenoma cells and about five when comparing adenoma to carcinoma cells. This list includes genes previously known to be associated with colon cancer as well as some novel genes. AVAILABILITY: R code for the empirical Bayes adjustment and step-down P-value calculation via resampling are available from the supplementary web-site. SUPPLEMENTARY INFORMATION: http://www.mathstat.gsu.edu/~matsnd/EB/supp.htm Susmita Datta, Glen A. Satten, Dale J. Benos, Jiazeng Xia, Martin J. Heslin, Somnath Datta |
Bioinform. | 2 |
| 2004 | Standardization and denoising algorithms for mass spectra to classify whole-organism bacterial specimensabstractMOTIVATION: Application of mass spectrometry in proteomics is a breakthrough in high-throughput analyses. Early applications have focused on protein expression profiles to differentiate among various types of tissue samples (e.g. normal versus tumor). Here our goal is to use mass spectra to differentiate bacterial species using whole-organism samples. The raw spectra are similar to spectra of tissue samples, raising some of the same statistical issues (e.g. non-uniform baselines and higher noise associated with higher baseline), but are substantially noisier. As a result, new preprocessing procedures are required before these spectra can be used for statistical classification. RESULTS: In this study, we introduce novel preprocessing steps that can be used with any mass spectra. These comprise a standardization step and a denoising step. The noise level for each spectrum is determined using only data from that spectrum. Only spectral features that exceed a threshold defined by the noise level are subsequently used for classification. Using this approach, we trained the Random Forest program to classify 240 mass spectra into four bacterial types. The method resulted in zero prediction errors in the training samples and in two test datasets having 240 and 300 spectra, respectively. Glen A. Satten, Somnath Datta, Hercules Moura, Adrian R. Woolfitt, Maria da G. Carvalho, George M. Carlone, Barun K. De, Antonis Pavlopoulos, John R. Barr 0002 |
Bioinform. | 1 |