Kathryn Roeder

dblp:19/10452 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-8869-6254ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Artificial intelligence and machine learning · 3
YearPublicationVenuePosition
2026 A framework to infer de novo exonic variants when parental genotypes are missing enhances association studies of autism
abstract
Abstract Motivation Gene-damaging mutations are highly informative for studies seeking to discover genes underlying developmental disorders. Traditionally, these de novo variants are recognized by evaluating high-quality DNA sequence from affected offspring and parents. However, when parental sequence is unavailable, methods are required to infer de novo status and use this inference for association studies. Results We use data from autism spectrum disorder to illustrate and evaluate methods. Separating de novo from rare inherited variants is challenging because the latter are far more common. Using a classifier for unbalanced data and variants of known inheritance class, we build an inheritance model and then a de novo score for variants when parental data are missing. Next, we propose a new Random Draw (RD) model to use this score for gene discovery. Built into an existing inferential framework, RD produces a more powerful gene-based association test and controls the false discovery rate. Availability and implementation Codes are available at Github (https://github.com/HaeunM/TADA-RD) and Zenodo (DOI: https://doi.org/10.5281/zenodo.18531769).
Haeun Moon, Laura Sloofman, Marina Natividad Avila, Lambertus Klei, Bernie Devlin, Joseph D. Buxbaum, Kathryn Roeder
Bioinform.7
2024 eSVD-DE: cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings
abstract
BACKGROUND: Single-cell RNA-sequencing (scRNA) datasets are becoming increasingly popular in clinical and cohort studies, but there is a lack of methods to investigate differentially expressed (DE) genes among such datasets with numerous individuals. While numerous methods exist to find DE genes for scRNA data from limited individuals, differential-expression testing for large cohorts of case and control individuals using scRNA data poses unique challenges due to substantial effects of human variation, i.e., individual-level confounding covariates that are difficult to account for in the presence of sparsely-observed genes. RESULTS: We develop the eSVD-DE, a matrix factorization that pools information across genes and removes confounding covariate effects, followed by a novel two-sample test in mean expression between case and control individuals. In general, differential testing after dimension reduction yields an inflation of Type-1 errors. However, we overcome this by testing for differences between the case and control individuals' posterior mean distributions via a hierarchical model. In previously published datasets of various biological systems, eSVD-DE has more accuracy and power compared to other DE methods typically repurposed for analyzing cohort-wide differential expression. CONCLUSIONS: eSVD-DE proposes a novel and powerful way to test for DE genes among cohorts after performing a dimension reduction. Accurate identification of differential expression on the individual level, instead of the cell level, is important for linking scRNA-seq studies to our understanding of the human population.
Kevin Z. Lin, Kathryn Roeder
BMC Bioinform.3
2021 An approach to gene-based testing accounting for dependence of tests among nearby genes
abstract
In genome-wide association studies (GWAS), it has become commonplace to test millions of single-nucleotide polymorphisms (SNPs) for phenotypic association. Gene-based testing can improve power to detect weak signal by reducing multiple testing and pooling signal strength. While such tests account for linkage disequilibrium (LD) structure of SNP alleles within each gene, current approaches do not capture LD of SNPs falling in different nearby genes, which can induce correlation of gene-based test statistics. We introduce an algorithm to account for this correlation. When a gene's test statistic is independent of others, it is assessed separately; when test statistics for nearby genes are strongly correlated, their SNPs are agglomerated and tested as a locus. To provide insight into SNPs and genes driving association within loci, we develop an interactive visualization tool to explore localized signal. We demonstrate our approach in the context of weakly powered GWAS for autism spectrum disorder, which is contrasted to more highly powered GWAS for schizophrenia and educational attainment. To increase power for these analyses, especially those for autism, we use adaptive $P$-value thresholding, guided by high-dimensional metadata modeled with gradient boosted trees, highlighting when and how it can be most useful. Notably our workflow is based on summary statistics.
Ronald Yurko, Kathryn Roeder, Bernie Devlin, Max G'Sell
Briefings Bioinform.2
2021 Identification of cell-type-specific marker genes from co-expression patterns in tissue samples
abstract
MOTIVATION: Marker genes, defined as genes that are expressed primarily in a single-cell type, can be identified from the single-cell transcriptome; however, such data are not always available for the many uses of marker genes, such as deconvolution of bulk tissue. Marker genes for a cell type, however, are highly correlated in bulk data, because their expression levels depend primarily on the proportion of that cell type in the samples. Therefore, when many tissue samples are analyzed, it is possible to identify these marker genes from the correlation pattern. RESULTS: To capitalize on this pattern, we develop a new algorithm to detect marker genes by combining published information about likely marker genes with bulk transcriptome data in the form of a semi-supervised algorithm. The algorithm then exploits the correlation structure of the bulk data to refine the published marker genes by adding or removing genes from the list. AVAILABILITY AND IMPLEMENTATION: We implement this method as an R package markerpen, hosted on CRAN (https://CRAN.R-project.org/package=markerpen). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiebiao Wang, Jing Lei 0003, Kathryn Roeder
Bioinform.4
2021 ESCO: single cell expression simulation incorporating gene co-expression
abstract
MOTIVATION: Gene-gene co-expression networks (GCN) are of biological interest for the useful information they provide for understanding gene-gene interactions. The advent of single cell RNA-sequencing allows us to examine more subtle gene co-expression occurring within a cell type. Many imputation and denoising methods have been developed to deal with the technical challenges observed in single cell data; meanwhile, several simulators have been developed for benchmarking and assessing these methods. Most of these simulators, however, either do not incorporate gene co-expression or generate co-expression in an inconvenient manner. RESULTS: Therefore, with the focus on gene co-expression, we propose a new simulator, ESCO, which adopts the idea of the copula to impose gene co-expression, while preserving the highlights of available simulators, which perform well for simulation of gene expression marginally. Using ESCO, we assess the performance of imputation methods on GCN recovery and find that imputation generally helps GCN recovery when the data are not too sparse, and the ensemble imputation method works best among leading methods. In contrast, imputation fails to help in the presence of an excessive fraction of zero counts, where simple data aggregating methods are a better choice. These findings are further verified with mouse and human brain cell data. AVAILABILITY AND IMPLEMENTATION: The ESCO implementation is available as R package ESCO. Users can either download the development version via github (https://github.com/JINJINT/ESCO) or the archived version via Zenodo (https://zenodo.org/record/4455890). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jinjin Tian, Jiebiao Wang, Kathryn Roeder
Bioinform.3
2020 Using multiple measurements of tissue to estimate subject- and cell-type-specific gene expression
abstract
MOTIVATION: Patterns of gene expression, quantified at the level of tissue or cells, can inform on etiology of disease. There are now rich resources for tissue-level (bulk) gene expression data, which have been collected from thousands of subjects, and resources involving single-cell RNA-sequencing (scRNA-seq) data are expanding rapidly. The latter yields cell type information, although the data can be noisy and typically are derived from a small number of subjects. RESULTS: Complementing these approaches, we develop a method to estimate subject- and cell-type-specific (CTS) gene expression from tissue using an empirical Bayes method that borrows information across multiple measurements of the same tissue per subject (e.g. multiple regions of the brain). Analyzing expression data from multiple brain regions from the Genotype-Tissue Expression project (GTEx) reveals CTS expression, which then permits downstream analyses, such as identification of CTS expression Quantitative Trait Loci (eQTL). AVAILABILITY AND IMPLEMENTATION: We implement this method as an R package MIND, hosted on https://github.com/randel/MIND. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiebiao Wang, Bernie Devlin, Kathryn Roeder
Bioinform.3
2015 MIRA: mutual information-based reporter algorithm for metabolic networks
abstract
doi: 10.1093/bioinformatics/btu290 Bioinformatics (2014) 30(12), i175–i184 The authors of the above article would like the following to be noted. Results reported for the reporter algorithm (RA) in the Results section are raw z-scores per metabolite, normalized by subtracting the sample mean and dividing by the sample standard deviation. Using the background correction as described in the original scoring scheme (Ideker et al., 2002; Patil and Nielsen, 2005) removes the hub metabolite bias we claim for the RA (for both normalization by k and k⁠). This affects only our comparisons with RA, and MIRA’s results are not changed. The background correction improves the empirical significance of the RA as shown in Figure 5A and C. After the background correction, RA is among the top three permutations. MIRA still performs better but with a smaller margin than shown in Figure 5. We have updated the lists of reporter metabolites (for RA) in the supplementary file.
A. Ercüment Çiçek, Kathryn Roeder, Gultekin Özsoyoglu
Bioinform.2
2014 MIRA: mutual information-based reporter algorithm for metabolic networks
abstract
MOTIVATION: Discovering the transcriptional regulatory architecture of the metabolism has been an important topic to understand the implications of transcriptional fluctuations on metabolism. The reporter algorithm (RA) was proposed to determine the hot spots in metabolic networks, around which transcriptional regulation is focused owing to a disease or a genetic perturbation. Using a z-score-based scoring scheme, RA calculates the average statistical change in the expression levels of genes that are neighbors to a target metabolite in the metabolic network. The RA approach has been used in numerous studies to analyze cellular responses to the downstream genetic changes. In this article, we propose a mutual information-based multivariate reporter algorithm (MIRA) with the goal of eliminating the following problems in detecting reporter metabolites: (i) conventional statistical methods suffer from small sample sizes, (ii) as z-score ranges from minus to plus infinity, calculating average scores can lead to canceling out opposite effects and (iii) analyzing genes one by one, then aggregating results can lead to information loss. MIRA is a multivariate and combinatorial algorithm that calculates the aggregate transcriptional response around a metabolite using mutual information. We show that MIRA's results are biologically sound, empirically significant and more reliable than RA. RESULTS: We apply MIRA to gene expression analysis of six knockout strains of Escherichia coli and show that MIRA captures the underlying metabolic dynamics of the switch from aerobic to anaerobic respiration. We also apply MIRA to an Autism Spectrum Disorder gene expression dataset. Results indicate that MIRA reports metabolites that highly overlap with recently found metabolic biomarkers in the autism literature. Overall, MIRA is a promising algorithm for detecting metabolic drug targets and understanding the relation between gene expression and metabolic activity. AVAILABILITY AND IMPLEMENTATION: The code is implemented in C# language using .NET framework. Project is available upon request.
A. Ercüment Çiçek, Kathryn Roeder, Gultekin Özsoyoglu
Bioinform.2
2012 Smooth-projected Neighborhood Pursuit for High-dimensional Nonparanormal Graph Estimation
abstract
Many statistical methods gain robustness and exibility by sacricing convenient computational structure. In this paper, we illustrate this fundamental tradeoff by studying a semiparametric graphical model estimation problem. We explain how new computational techniques help to solve this type of problem. In particularly, we propose a smooth-projected neighborhood pursuit method for efciently estimating high dimensional nonparanormal graphs with theoretical guarantees. Besides new computational and theoretical analysis, we also provide an alternative view to analyze the tradeoff between computational efciency and statistical error under a smoothing optimization framework. We also report experimental results on text and stock datasets.
Tuo Zhao, Kathryn Roeder, Han Liu 0001
NIPS2
2012 The huge Package for High-dimensional Undirected Graph Estimation in R
Tuo Zhao, Han Liu 0001, Kathryn Roeder, John D. Lafferty, Larry A. Wasserman
J. Mach. Learn. Res.3
2010 Stability Approach to Regularization Selection (StARS) for High Dimensional Graphical Models
abstract
A challenging problem in estimating high-dimensional graphical models is to choose the regularization parameter in a data-dependent way. The standard techniques include $K$-fold cross-validation ($K$-CV), Akaike information criterion (AIC), and Bayesian information criterion (BIC). Though these methods work well for low-dimensional problems, they are not suitable in high dimensional settings. In this paper, we present StARS: a new stability-based method for choosing the regularization parameter in high dimensional inference for undirected graphs. The method has a clear interpretation: we use the least amount of regularization that simultaneously makes a graph sparse and replicable under random sampling. This interpretation requires essentially no conditions. Under mild conditions, we show that StARS is partially sparsistent in terms of graph estimation: i.e. with high probability, all the true edges will be included in the selected model even when the graph size asymptotically increases with the sample size. Empirically, the performance of StARS is compared with the state-of-the-art model selection procedures, including $K$-CV, AIC, and BIC, on both synthetic data and a real microarray dataset. StARS outperforms all competing procedures.
Han Liu 0001, Kathryn Roeder, Larry A. Wasserman
NIPS2