EDBT 2026 Demo / reviewers in the wild / expert
Andrew S. Allen
dblp:63/1436
· DBLP profile ↗
14ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-7232-2143ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genomics |
1.0 | 3 | 2025 | Bayesian estimation of genetic regulatory effects in high-throughput reporter assays · Bioinform. 2020 High-throughput interpretation of gene structure changes in human and nonhuman resequencing data, using ACE · Bioinform. 2017 Bayesian estimation of allele-specific expression in the presence of phasing uncertainty · Bioinform. 2025 |
Bioinformatics and computational biology › gene expression analysis › gene expression quantification
allele-specific expression |
0.9 | 1 | 2025 | Bayesian estimation of allele-specific expression in the presence of phasing uncertainty · Bioinform. 2025 |
Bioinformatics and computational biology
cancer genomics |
0.5 | 1 | 2021 | A Bayesian hierarchical model to estimate DNA methylation conservation in colorectal tumors · Bioinform. 2021 |
Bioinformatics and computational biology
epigenomics |
0.5 | 1 | 2021 | A Bayesian hierarchical model to estimate DNA methylation conservation in colorectal tumors · Bioinform. 2021 |
Bioinformatics and computational biology › gene regulation › regulatory variant interpretation
regulatory variant effect prediction |
0.4 | 1 | 2020 | Bayesian estimation of genetic regulatory effects in high-throughput reporter assays · Bioinform. 2020 |
Bioinformatics and computational biology › drug discovery
high-throughput screening |
0.3 | 1 | 2018 | bcSeq: an R package for fast sequence mapping in high-throughput shRNA and CRISPR screens · Bioinform. 2018 |
Bioinformatics and computational biology › sequence alignment
sequence mapping |
0.3 | 1 | 2018 | bcSeq: an R package for fast sequence mapping in high-throughput shRNA and CRISPR screens · Bioinform. 2018 |
Bioinformatics and computational biology › gene regulation
regulatory genomics |
0.3 | 1 | 2017 | Quantifying the Impact of Non-coding Variants on Transcription Factor-DNA Binding · RECOMB 2017 |
Bioinformatics and computational biology › gene regulation
transcription factor binding |
0.3 | 1 | 2017 | Quantifying the Impact of Non-coding Variants on Transcription Factor-DNA Binding · RECOMB 2017 |
Bioinformatics and computational biology › clinical bioinformatics
variant interpretation |
0.3 | 1 | 2017 | High-throughput interpretation of gene structure changes in human and nonhuman resequencing data, using ACE · Bioinform. 2017 |
Bioinformatics and computational biology › statistical genetics
genotype-phenotype analysis |
0.3 | 1 | 2025 | Bayesian estimation of allele-specific expression in the presence of phasing uncertainty · Bioinform. 2025 |
Bioinformatics and computational biology › genomics › variant annotation
non-coding variant functional annotation |
0.1 | 1 | 2020 | Bayesian estimation of genetic regulatory effects in high-throughput reporter assays · Bioinform. 2020 |
Bioinformatics and computational biology › genome annotation
genomic variant annotation |
0.1 | 1 | 2011 | SVA: software for annotating and visualizing sequenced human genomes · Bioinform. 2011 |
Bioinformatics and computational biology › genomics › genome visualization
variant visualization |
0.1 | 1 | 2011 | SVA: software for annotating and visualizing sequenced human genomes · Bioinform. 2011 |
Bioinformatics and computational biology › sequence analysis
sequencing error modeling |
0.1 | 1 | 2018 | bcSeq: an R package for fast sequence mapping in high-throughput shRNA and CRISPR screens · Bioinform. 2018 |
Bioinformatics and computational biology › transcriptomics › transcriptome annotation
transcript annotation |
0.1 | 1 | 2017 | High-throughput interpretation of gene structure changes in human and nonhuman resequencing data, using ACE · Bioinform. 2017 |
Methods — techniques the papers use, named apart from their topics
bayesian hierarchical model · 1.8empirical phasing error rates · 0.9stan · 0.5markov chain monte carlo · 0.5poisson-binomial distribution · 0.4trie data structure · 0.3statistical sequencing error model · 0.3quantitative modeling · 0.3haplotype reconstruction · 0.3RNA-seq analysis · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bayesian estimation of allele-specific expression in the presence of phasing uncertaintyabstractMOTIVATION: Allele-specific expression (ASE) analyses aim to detect imbalanced expression of maternal versus paternal copies of an autosomal gene. Such allelic imbalance can result from a variety of cis-acting causes, including disruptive mutations within one copy of a gene that impact the stability of transcripts, as well as regulatory variants outside the gene that impact transcription initiation. Current methods for ASE estimation suffer from a number of shortcomings, such as relying on only one variant within a gene, assuming perfect phasing information across multiple variants within a gene, or failing to account for alignment biases and possible genotyping errors. RESULTS: We developed BEASTIE, a Bayesian hierarchical model designed for precise ASE quantification at the gene level, based on given genotypes and RNA-Seq data. BEASTIE addresses the complexities of allelic mapping bias, genotyping error, and phasing errors by incorporating empirical phasing error rates derived from Genome-in-a-Bottle individual NA12878. BEASTIE surpasses existing methods in accuracy, especially in scenarios with high phasing errors. This improvement is critical for identifying rare genetic variants often obscured by such errors. Through rigorous validation on simulated data and application to real data from the 1000 Genomes Project, we establish the robustness of BEASTIE. These findings underscore the value of BEASTIE in revealing patterns of ASE across gene sets and pathways. AVAILABILITY AND IMPLEMENTATION: The software is freely available from Github (https://github.com/x811zou/BEASTIE); and Zendo (DOI: 10.5281/zenodo.15062124). Xue Zou, Zachary W. Gomez, Timothy E. Reddy, Andrew S. Allen, William H. Majoros |
Bioinform. | 4 |
| 2022 | Focused goodness of fit tests for gene set analysesabstractGene set-based signal detection analyses are used to detect an association between a trait and a set of genes by accumulating signals across the genes in the gene set. Since signal detection is concerned with identifying whether any of the genes in the gene set are non-null, a goodness-of-fit (GOF) test can be used to compare whether the observed distribution of gene-level tests within the gene set agrees with the theoretical null distribution. Here, we present a flexible gene set-based signal detection framework based on tail-focused GOF statistics. We show that the power of the various statistics in this framework depends critically on two parameters: the proportion of genes within the gene set that are non-null and the degree of separation between the null and alternative distributions of the gene-level tests. We give guidance on which statistic to choose for a given situation and implement the methods in a fast and user-friendly R package, wHC (https://github.com/mqzhanglab/wHC). Finally, we apply these methods to a whole exome sequencing study of amyotrophic lateral sclerosis. Sahar Gelfman, Cristiane Araujo Martins Moreno, Janice M. McCarthy, Matthew B. Harms, David B. Goldstein, Andrew S. Allen |
Briefings Bioinform. | 7 |
| 2021 | A Bayesian hierarchical model to estimate DNA methylation conservation in colorectal tumorsabstractMOTIVATION: Conservation is broadly used to identify biologically important (epi)genomic regions. In the case of tumor growth, preferential conservation of DNA methylation can be used to identify areas of particular functional importance to the tumor. However, reliable assessment of methylation conservation based on multiple tissue samples per patient requires the decomposition of methylation variation at multiple levels. RESULTS: We developed a Bayesian hierarchical model that allows for variance decomposition of methylation on three levels: between-patient normal tissue variation, between-patient tumor-effect variation and within-patient tumor variation. We then defined a model-based conservation score to identify loci of reduced within-tumor methylation variation relative to between-patient variation. We fit the model to multi-sample methylation array data from 21 colorectal cancer (CRC) patients using a Monte Carlo Markov Chain algorithm (Stan). Sets of genes implicated in CRC tumorigenesis exhibited preferential conservation, demonstrating the model's ability to identify functionally relevant genes based on methylation conservation. A pathway analysis of preferentially conserved genes implicated several CRC relevant pathways and pathways related to neoantigen presentation and immune evasion. Our findings suggest that preferential methylation conservation may be used to identify novel gene targets that are not consistently mutated in CRC. The flexible structure makes the model amenable to the analysis of more complex multi-sample data structures. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in the NCBI GEO Database, under accession code GSE166212. The R analysis code is available at https://github.com/kevin-murgas/DNAmethylation-hierarchicalmodel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kevin A. Murgas, Yanlin Ma, Lidea Shahidi, Sayan Mukherjee 0001, Andrew S. Allen, Darryl Shibata, Marc D. Ryser |
Bioinform. | 5 |
| 2020 | Bayesian estimation of genetic regulatory effects in high-throughput reporter assaysabstractMOTIVATION: High-throughput reporter assays dramatically improve our ability to assign function to noncoding genetic variants, by measuring allelic effects on gene expression in the controlled setting of a reporter gene. Unlike genetic association tests, such assays are not confounded by linkage disequilibrium when loci are independently assayed. These methods can thus improve the identification of causal disease mutations. While work continues on improving experimental aspects of these assays, less effort has gone into developing methods for assessing the statistical significance of assay results, particularly in the case of rare variants captured from patient DNA. RESULTS: We describe a Bayesian hierarchical model, called Bayesian Inference of Regulatory Differences, which integrates prior information and explicitly accounts for variability between experimental replicates. The model produces substantially more accurate predictions than existing methods when allele frequencies are low, which is of clear advantage in the search for disease-causing variants in DNA captured from patient cohorts. Using the model, we demonstrate a clear tradeoff between variant sequencing coverage and numbers of biological replicates, and we show that the use of additional biological replicates decreases variance in estimates of effect size, due to the properties of the Poisson-binomial distribution. We also provide a power and sample size calculator, which facilitates decision making in experimental design parameters. AVAILABILITY AND IMPLEMENTATION: The software is freely available from www.geneprediction.org/bird. The experimental design web tool can be accessed at http://67.159.92.22:8080. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. William H. Majoros, Young-Sook Kim, Alejandro Barrera, Fan Li 0021, Xingyan Wang, Sarah J. Cunningham, Graham D. Johnson, William L. Lowe, Denise M. Scholtens, M. Geoffrey Hayes, Timothy E. Reddy, Andrew S. Allen |
Bioinform. | 13 |
| 2019 | Efficient estimation of grouped survival modelsabstractBACKGROUND: Time- and dose-to-event phenotypes used in basic science and translational studies are commonly measured imprecisely or incompletely due to limitations of the experimental design or data collection schema. For example, drug-induced toxicities are not reported by the actual time or dose triggering the event, but rather are inferred from the cycle or dose to which the event is attributed. This exemplifies a prevalent type of imprecise measurement called grouped failure time, where times or doses are restricted to discrete increments. Failure to appropriately account for the grouped nature of the data, when present, may lead to biased analyses. RESULTS: We present groupedSurv, an R package which implements a statistically rigorous and computationally efficient approach for conducting genome-wide analyses based on grouped failure time phenotypes. Our approach accommodates adjustments for baseline covariates, and analysis at the variant or gene level. We illustrate the statistical properties of the approach and computational performance of the package by simulation. We present the results of a reanalysis of a published genome-wide study to identify common germline variants associated with the risk of taxane-induced peripheral neuropathy in breast cancer patients. CONCLUSIONS: groupedSurv enables fast and rigorous genome-wide analysis on the basis of grouped failure time phenotypes at the variant, gene or pathway level. The package is freely available under a public license through the Comprehensive R Archive Network. Jiaxing Lin, Alexander B. Sibley, Tracy Truong, Katherina C. Chua, Janice M. McCarthy, Deanna L. Kroetz, Andrew S. Allen, Kouros Owzar |
BMC Bioinform. | 9 |
| 2018 | bcSeq: an R package for fast sequence mapping in high-throughput shRNA and CRISPR screensabstractSummary: CRISPR-Cas9 and shRNA high-throughput sequencing screens have abundant applications for basic and translational research. Methods and tools for the analysis of these screens must properly account for sequencing error, resolve ambiguous mappings among similar sequences in the barcode library in a statistically principled manner, and be computationally efficient. Herein we present bcSeq, an open source R package that implements a fast and parallelized algorithm for mapping high-throughput sequencing reads to a barcode library while tolerating sequencing error. The algorithm uses a Trie data structure for speed and resolves ambiguous mappings by using a statistical sequencing error model based on Phred scores for each read. Availability and implementation: The package source code and an accompanying tutorial are available at http://bioconductor.org/packages/bcSeq/. Supplementary information: Supplementary data are available at Bioinformatics online. Jiaxing Lin, Jeremy Gresham, Tongrong Wang, So Young Kim, James Alvarez, Jeffrey S. Damrauer, Scott Floyd, Joshua A. Granek, Andrew S. Allen, Cliburn Chan, Jichun Xie, Kouros Owzar |
Bioinform. | 9 |
| 2018 | meaRtools: An R package for the analysis of neuronal networks recorded on microelectrode arraysabstractHere we present an open-source R package 'meaRtools' that provides a platform for analyzing neuronal networks recorded on Microelectrode Arrays (MEAs). Cultured neuronal networks monitored with MEAs are now being widely used to characterize in vitro models of neurological disorders and to evaluate pharmaceutical compounds. meaRtools provides core algorithms for MEA spike train analysis, feature extraction, statistical analysis and plotting of multiple MEA recordings with multiple genotypes and treatments. meaRtools functionality covers novel solutions for spike train analysis, including algorithms to assess electrode cross-correlation using the spike train tiling coefficient (STTC), mutual information, synchronized bursts and entropy within cultured wells. Also integrated is a solution to account for bursts variability originating from mixed-cell neuronal cultures. The package provides a statistical platform built specifically for MEA data that can combine multiple MEA recordings and compare extracted features between different genetic models or treatments. We demonstrate the utilization of meaRtools to successfully identify epilepsy-like phenotypes in neuronal networks from Celf4 knockout mice. The package is freely available under the GPL license (GPL> = 3) and is updated frequently on the CRAN web-server repository. The package, along with full documentation can be downloaded from: https://cran.r-project.org/web/packages/meaRtools/. Sahar Gelfman, Quanli Wang, Yi-Fan Lu, Diana Hall, Christopher D. Bostick, Ryan Dhindsa, Matt Halvorsen, K. Melodi McSweeney, Ellese Cotterill, Tom Edinburgh, Michael A. Beaumont, Wayne N. Frankel, Slavé Petrovski, Andrew S. Allen, Michael J. Boland, David B. Goldstein, Stephen J. Eglen |
PLoS Comput. Biol. | 14 |
| 2017 | Quantifying the Impact of Non-coding Variants on Transcription Factor-DNA Binding
Jingkang Zhao, Dongshunyi Li, Jungkyun Seo, Andrew S. Allen, Raluca Gordân |
RECOMB | 4 |
| 2017 | High-throughput interpretation of gene structure changes in human and nonhuman resequencing data, using ACEabstractMOTIVATION: The accurate interpretation of genetic variants is critical for characterizing genotype-phenotype associations. Because the effects of genetic variants can depend strongly on their local genomic context, accurate genome annotations are essential. Furthermore, as some variants have the potential to disrupt or alter gene structure, variant interpretation efforts stand to gain from the use of individualized annotations that account for differences in gene structure between individuals or strains. RESULTS: We describe a suite of software tools for identifying possible functional changes in gene structure that may result from sequence variants. ACE ('Assessing Changes to Exons') converts phased genotype calls to a collection of explicit haplotype sequences, maps transcript annotations onto them, detects gene-structure changes and their possible repercussions, and identifies several classes of possible loss of function. Novel transcripts predicted by ACE are commonly supported by spliced RNA-seq reads, and can be used to improve read alignment and transcript quantification when an individual-specific genome sequence is available. Using publicly available RNA-seq data, we show that ACE predictions confirm earlier results regarding the quantitative effects of nonsense-mediated decay, and we show that predicted loss-of-function events are highly concordant with patterns of intolerance to mutations across the human population. ACE can be readily applied to diverse species including animals and plants, making it a broadly useful tool for use in eukaryotic population-based resequencing projects, particularly for assessing the joint impact of all variants at a locus. AVAILABILITY AND IMPLEMENTATION: ACE is written in open-source C ++ and Perl and is available from geneprediction.org/ACE. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online. William H. Majoros, Michael Campbell, Carson Holt, Erin K. DeNardo, Doreen Ware, Andrew S. Allen, Mark Yandell, Timothy E. Reddy |
Bioinform. | 6 |
| 2017 | Mapping eQTL by leveraging multiple tissues and DNA methylationabstractBACKGROUND: DNA methylation is an important tissue-specific epigenetic event that influences transcriptional regulation of gene expression. Differentially methylated CpG sites may act as mediators between genetic variation and gene expression, and this relationship can be exploited while mapping multi-tissue expression quantitative trait loci (eQTL). Current multi-tissue eQTL mapping techniques are limited to only exploiting gene expression patterns across multiple tissues either in a joint tissue or tissue-by-tissue frameworks. We present a new statistical approach that enables us to model the effect of germ-line variation on tissue-specific gene expression in the presence of effects due to DNA methylation. RESULTS: Our method efficiently models genetic and epigenetic variation to identify genomic regions of interest containing combinations of mRNA transcripts, CpG sites, and SNPs by jointly testing for genotypic effect and higher order interaction effects between genotype, methylation and tissues. We demonstrate using Monte Carlo simulations that our approach, in the presence of both genetic and DNA methylation effects, gives an improved performance (in terms of statistical power) to detect eQTLs over the current eQTL mapping approaches. When applied to an array-based dataset from 150 neuropathologically normal adult human brains, our method identifies eQTLs that were undetected using standard tissue-by-tissue or joint tissue eQTL mapping techniques. As an example, our method identifies eQTLs by leveraging methylated CpG sites in a LIM homeobox member gene (LHX9), which may have a role in the neural development. CONCLUSIONS: Our score test-based approach does not need parameter estimation under the alternative hypothesis. As a result, our model parameters are estimated only once for each mRNA - CpG pair. Our model specifically studies the effects of non-coding regions of DNA (in this case, CpG sites) on mapping eQTLs. However, we can easily model micro-RNAs instead of CpG sites to study the effects of post-transcriptional events in mapping eQTL. Our model's flexible framework also allows us to investigate other genomic events such as alternative gene splicing by extending our model to include gene isoform-specific data. Chaitanya R. Acharya, Kouros Owzar, Andrew S. Allen |
BMC Bioinform. | 3 |
| 2016 | Exploiting expression patterns across multiple tissues to map expression quantitative trait lociabstractBACKGROUND: In order to better understand complex diseases, it is important to understand how genetic variation in the regulatory regions affects gene expression. Genetic variants found in these regulatory regions have been shown to activate transcription in a tissue-specific manner. Therefore, it is important to map the aforementioned expression quantitative trait loci (eQTL) using a statistically disciplined approach that jointly models all the tissues and makes use of all the information available to maximize the power of eQTL mapping. In this context, we are proposing a score test-based approach where we model tissue-specificity as a random effect and investigate an overall shift in the gene expression combined with tissue-specific effects due to genetic variants. RESULTS: Our approach has 1) a distinct computational edge, and 2) comparable performance in terms of statistical power over other currently existing joint modeling approaches such as MetaTissue eQTL and eQTL-BMA. Using simulations, we show that our method increases the power to detect eQTLs when compared to a tissue-by-tissue approach and can exceed the performance, in terms of computational speed, of MetaTissue eQTL and eQTL-BMA. We apply our method to two publicly available expression datasets from normal human brains, one comprised of four brain regions from 150 neuropathologically normal samples and another comprised of ten brain regions from 134 neuropathologically normal samples, and show that by using our method and jointly analyzing multiple brain regions, we identify eQTLs within more genes when compared to three often used existing methods. CONCLUSIONS: Since we employ a score test-based approach, there is no need for parameter estimation under the alternative hypothesis. As a result, model parameters only have to be estimated once per genome, significantly decreasing computation time. Our method also accommodates the analysis of next- generation sequencing data. As an example, by modeling gene transcripts in an analogous fashion to tissues in our current formulation one would be able to test for both a variant overall effect across all isoforms of a gene as well as transcript-specific effects. We implement our approach within the R package JAGUAR, which is now available at the Comprehensive R Archive Network repository. Chaitanya R. Acharya, Janice M. McCarthy, Kouros Owzar, Andrew S. Allen |
BMC Bioinform. | 4 |
| 2013 | Leveraging Prior Information to Detect Causal Variants via Multi-Variant RegressionabstractAlthough many methods are available to test sequence variants for association with complex diseases and traits, methods that specifically seek to identify causal variants are less developed. Here we develop and evaluate a Bayesian hierarchical regression method that incorporates prior information on the likelihood of variant causality through weighting of variant effects. By simulation studies using both simulated and real sequence variants, we compared a standard single variant test for analyzing variant-disease association with the proposed method using different weighting schemes. We found that by leveraging linkage disequilibrium of variants with known GWAS signals and sequence conservation (phastCons), the proposed method provides a powerful approach for detecting causal variants while controlling false positives. Nanye Long, Samuel P. Dickson, Jessica M. Maia, Hee Shin Kim, Andrew S. Allen |
PLoS Comput. Biol. | 6 |
| 2011 | SVA: software for annotating and visualizing sequenced human genomesabstractSUMMARY: Here we present Sequence Variant Analyzer (SVA), a software tool that assigns a predicted biological function to variants identified in next-generation sequencing studies and provides a browser to visualize the variants in their genomic contexts. SVA also provides for flexible interaction with software implementing variant association tests allowing users to consider both the bioinformatic annotation of identified variants and the strength of their associations with studied traits. We illustrate the annotation features of SVA using two simple examples of sequenced genomes that harbor Mendelian mutations. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://www.svaproject.org. Dongliang Ge, Elizabeth K. Ruzzo, Kevin V. Shianna, Kimberly Pelak, Erin L. Heinzen, Anna C. Need, Elizabeth T. Cirulli, Jessica M. Maia, Samuel P. Dickson, Mingfu Zhu, Abanish Singh, Andrew S. Allen, David B. Goldstein |
Bioinform. | 13 |
| 2008 | Invited Keynote Talk: Haplotype Sharing for Genome-Wide Case-Control Association Studies
Andrew S. Allen |
ISBRA | 1 |