VLDB 2026 Research / reviewers in the wild / expert
Alison A. Motsinger-Reif
dblp:m/AlisonAMotsinger · also Alison A. Motsinger, Alison Anne Motsinger-Reif
· DBLP profile ↗
15ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-1346-2493ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient gene set analysis for DNA methylation addressing probe dependency and biasabstractMOTIVATION: Gene Set Enrichment Analysis (GSEA) is widely used to interpret DNA methylation data by associating differentially methylated sites with biological pathways. However, existing GSEA methods struggle with several challenges in methylation data, including probe dependency, probe number bias, and the complexity of gene-probe mapping. These limitations can lead to biased enrichment results, reduced statistical power, and computational inefficiencies. RESULTS: We introduce gsGene and gsPG, two novel GSEA methods specifically designed for DNA methylation data. gsGene aggregates association signals at gene level while correcting for probe dependency and probe number bias, enabling more biologically meaningful enrichment analysis. gsPG takes a different approach by conducting gene set enrichment using summary statistics for independent probe groups based on gene annotation, mitigating biases from multi-mapping probes. Both methods improve computational efficiency, enhance statistical power, and effectively control type I error rates. Comprehensive evaluations in two large datasets demonstrate superior performance compared to existing methods. Furthermore, we propose a novel beta distribution fitting strategy to improve enrichment P-value estimation, providing a computationally efficient alternative to traditional permutation-based gene set methods. AVAILABILITY AND IMPLEMENTATION: These methods are implemented in the R package dmGsea, which is freely available on GitHub and Bioconductor (DOI: 10.18129/B9.bioc.dmGsea). The package supports Illumina 450K, EPIC, and mouse methylation arrays and can be extended to other omics data with user-provided probe-to-gene mapping annotations. Zongli Xu, Alison A. Motsinger-Reif, Liang Niu |
Bioinform. | 2 |
| 2024 | Model-based evaluation of spatiotemporal data reduction methods with unknown ground truth through optimal visualization and interpretability metricsabstractOptimizing and benchmarking data reduction methods for dynamic or spatial visualization and interpretation (DSVI) face challenges due to many factors, including data complexity, lack of ground truth, time-dependent metrics, dimensionality bias and different visual mappings of the same data. Current studies often focus on independent static visualization or interpretability metrics that require ground truth. To overcome this limitation, we propose the MIBCOVIS framework, a comprehensive and interpretable benchmarking and computational approach. MIBCOVIS enhances the visualization and interpretability of high-dimensional data without relying on ground truth by integrating five robust metrics, including a novel time-ordered Markov-based structural metric, into a semi-supervised hierarchical Bayesian model. The framework assesses method accuracy and considers interaction effects among metric features. We apply MIBCOVIS using linear and nonlinear dimensionality reduction methods to evaluate optimal DSVI for four distinct dynamic and spatial biological processes captured by three single-cell data modalities: CyTOF, scRNA-seq and CODEX. These data vary in complexity based on feature dimensionality, unknown cell types and dynamic or spatial differences. Unlike traditional single-summary score approaches, MIBCOVIS compares accuracy distributions across methods. Our findings underscore the joint evaluation of visualization and interpretability, rather than relying on separate metrics. We reveal that prioritizing average performance can obscure method feature performance. Additionally, we explore the impact of data complexity on visualization and interpretability. Specifically, we provide optimal parameters and features and recommend methods, like the optimized variational contractive autoencoder, for targeted DSVI for various data complexities. MIBCOVIS shows promise for evaluating dynamic single-cell atlases and spatiotemporal data reduction models. Komlan Atitey, Alison A. Motsinger-Reif, Benedict Anchang |
Briefings Bioinform. | 2 |
| 2024 | Taxanorm: a novel taxa-specific normalization approach for microbiome dataabstractBACKGROUND: In high-throughput sequencing studies, sequencing depth, which quantifies the total number of reads, varies across samples. Unequal sequencing depth can obscure true biological signals of interest and prevent direct comparisons between samples. To remove variability due to differential sequencing depth, taxa counts are usually normalized before downstream analysis. However, most existing normalization methods scale counts using size factors that are sample specific but not taxa specific, which can result in over- or under-correction for some taxa. RESULTS: We developed TaxaNorm, a novel normalization method based on a zero-inflated negative binomial model. This method assumes the effects of sequencing depth on mean and dispersion vary across taxa. Incorporating the zero-inflation part can better capture the nature of microbiome data. We also propose two corresponding diagnosis tests on the varying sequencing depth effect for validation. We find that TaxaNorm achieves comparable performance to existing methods in most simulation scenarios in downstream analysis and reaches a higher power for some cases. Specifically, it balances power and false discovery control well. When applying the method in a real dataset, TaxaNorm has improved performance when correcting technical bias. CONCLUSION: TaxaNorm both sample- and taxon- specific bias by introducing an appropriate regression framework in the microbiome data, which aids in data interpretation and visualization. The 'TaxaNorm' R package is freely available through the CRAN repository https://CRAN.R-project.org/package=TaxaNorm and the source code can be downloaded at https://github.com/wangziyue57/TaxaNorm . Dillon Lloyd, Alison A. Motsinger-Reif |
BMC Bioinform. | 4 |
| 2022 | Visualization, benchmarking and characterization of nested single-cell heterogeneity as dynamic forest mixturesabstractA major topic of debate in developmental biology centers on whether development is continuous, discontinuous, or a mixture of both. Pseudo-time trajectory models, optimal for visualizing cellular progression, model cell transitions as continuous state manifolds and do not explicitly model real-time, complex, heterogeneous systems and are challenging for benchmarking with temporal models. We present a data-driven framework that addresses these limitations with temporal single-cell data collected at discrete time points as inputs and a mixture of dependent minimum spanning trees (MSTs) as outputs, denoted as dynamic spanning forest mixtures (DSFMix). DSFMix uses decision-tree models to select genes that account for variations in multimodality, skewness and time. The genes are subsequently used to build the forest using tree agglomerative hierarchical clustering and dynamic branch cutting. We first motivate the use of forest-based algorithms compared to single-tree approaches for visualizing and characterizing developmental processes. We next benchmark DSFMix to pseudo-time and temporal approaches in terms of feature selection, time correlation, and network similarity. Finally, we demonstrate how DSFMix can be used to visualize, compare and characterize complex relationships during biological processes such as epithelial-mesenchymal transition, spermatogenesis, stem cell pluripotency, early transcriptional response from hormones and immune response to coronavirus disease. Our results indicate that the expression of genes during normal development exhibits a high proportion of non-uniformly distributed profiles that are mostly right-skewed and multimodal; the latter being a characteristic of major steady states during development. Our study also identifies and validates gene signatures driving complex dynamic processes during somatic or germline differentiation. Benedict Anchang, Raul Mendez-Giraldez, Xiaojiang Xu, Trevor K. Archer, Guang Hu 0005, Sylvia K. Plevritis, Alison A. Motsinger-Reif, Jian-Liang Li |
Briefings Bioinform. | 8 |
| 2021 | Knockoff boosted tree for model-free variable selectionabstractMOTIVATION: The recently proposed knockoff filter is a general framework for controlling the false discovery rate (FDR) when performing variable selection. This powerful new approach generates a 'knockoff' of each variable tested for exact FDR control. Imitation variables that mimic the correlation structure found within the original variables serve as negative controls for statistical inference. Current applications of knockoff methods use linear regression models and conduct variable selection only for variables existing in model functions. Here, we extend the use of knockoffs for machine learning with boosted trees, which are successful and widely used in problems where no prior knowledge of model function is required. However, currently available importance scores in tree models are insufficient for variable selection with FDR control. RESULTS: We propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. We extend the current knockoff method to model-free variable selection through the use of tree-based models. Additionally, we propose and evaluate two new sampling methods for generating knockoffs, namely the sparse covariance and principal component knockoff methods. We test and compare these methods with the original knockoff method regarding their ability to control type I errors and power. In simulation tests, we compare the properties and performance of importance test statistics of tree models. The results include different combinations of knockoffs and importance test statistics. We consider scenarios that include main-effect, interaction, exponential and second-order models while assuming the true model structures are unknown. We apply our algorithm for tumor purity estimation and tumor classification using Cancer Genome Atlas (TCGA) gene expression data. Our results show improved discrimination between difficult-to-discriminate cancer types. AVAILABILITY AND IMPLEMENTATION: The proposed algorithm is included in the KOBT package, which is available at https://cran.r-project.org/web/packages/KOBT/index.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alison A. Motsinger-Reif |
Bioinform. | 3 |
| 2019 | Identifying individual risk rare variants using protein structure guided local tests (POINT)abstractRare variants are of increasing interest to genetic association studies because of their etiological contributions to human complex diseases. Due to the rarity of the mutant events, rare variants are routinely analyzed on an aggregate level. While aggregation analyses improve the detection of global-level signal, they are not able to pinpoint causal variants within a variant set. To perform inference on a localized level, additional information, e.g., biological annotation, is often needed to boost the information content of a rare variant. Following the observation that important variants are likely to cluster together on functional domains, we propose a protein structure guided local test (POINT) to provide variant-specific association information using structure-guided aggregation of signal. Constructed under a kernel machine framework, POINT performs local association testing by borrowing information from neighboring variants in the 3-dimensional protein space in a data-adaptive fashion. Besides merely providing a list of promising variants, POINT assigns each variant a p-value to permit variant ranking and prioritization. We assess the selection performance of POINT using simulations and illustrate how it can be used to prioritize individual rare variants in PCSK9, ANGPTL4 and CETP in the Action to Control Cardiovascular Risk in Diabetes (ACCORD) clinical trial data. Rachel Marceau West, Wenbin Lu, Daniel M. Rotroff, Mélaine A. Kuenemann, Sheng-Mao Chang, Michael C. Wu, Michael J. Wagner, John B. Buse, Alison A. Motsinger-Reif, Denis Fourches, Jung-Ying Tzeng |
PLoS Comput. Biol. | 9 |
| 2014 | Bayesian neural networks for detecting epistasis in genetic association studiesabstractBACKGROUND: Discovering causal genetic variants from large genetic association studies poses many difficult challenges. Assessing which genetic markers are involved in determining trait status is a computationally demanding task, especially in the presence of gene-gene interactions. RESULTS: A non-parametric Bayesian approach in the form of a Bayesian neural network is proposed for use in analyzing genetic association studies. Demonstrations on synthetic and real data reveal they are able to efficiently and accurately determine which variants are involved in determining case-control status. By using graphics processing units (GPUs) the time needed to build these models is decreased by several orders of magnitude. In comparison with commonly used approaches for detecting interactions, Bayesian neural networks perform very well across a broad spectrum of possible genetic relationships. CONCLUSIONS: The proposed framework is shown to be a powerful method for detecting causal SNPs while being computationally efficient enough to handle large datasets. Andrew L. Beam, Alison A. Motsinger-Reif, Jon Doyle |
BMC Bioinform. | 2 |
| 2011 | The power of quantitative grammatical evolution neural networks to detect gene-gene interactionsabstractApplying grammatical evolution to evolve neural networks (GENN) has been increasing used in genetic epidemiology to detect gene-gene or gene-environment interactions, also known as epistasis, in high dimensional data. GENN approaches have previously been shown to be highly successful in a range of simulated and real case-control studies, and has recently been applied to quantitative traits. In the current study, we evaluate the potential of an application of GENN to quantitative traits (QTGENN) to a range of simulated genetic models. We demonstrate the power of the approach, and compare this power to more traditional linear regression analysis approaches. We find that the QTGENN approach has relatively high power to detect both single-locus models as well as several completely epistatic two-locus models, and favorably compares to the regression methods. Nicholas E. Hardison, Alison A. Motsinger-Reif |
GECCO | 2 |
| 2010 | A comparison of internal validation techniques for multifactor dimensionality reductionabstractBACKGROUND: It is hypothesized that common, complex diseases may be due to complex interactions between genetic and environmental factors, which are difficult to detect in high-dimensional data using traditional statistical approaches. Multifactor Dimensionality Reduction (MDR) is the most commonly used data-mining method to detect epistatic interactions. In all data-mining methods, it is important to consider internal validation procedures to obtain prediction estimates to prevent model over-fitting and reduce potential false positive findings. Currently, MDR utilizes cross-validation for internal validation. In this study, we incorporate the use of a three-way split (3WS) of the data in combination with a post-hoc pruning procedure as an alternative to cross-validation for internal model validation to reduce computation time without impairing performance. We compare the power to detect true disease causing loci using MDR with both 5- and 10-fold cross-validation to MDR with 3WS for a range of single-locus and epistatic disease models. Additionally, we analyze a dataset in HIV immunogenetics to demonstrate the results of the two strategies on real data. RESULTS: MDR with 3WS is computationally approximately five times faster than 5-fold cross-validation. The power to find the exact true disease loci without detecting false positive loci is higher with 5-fold cross-validation than with 3WS before pruning. However, the power to find the true disease causing loci in addition to false positive loci is equivalent to the 3WS. With the incorporation of a pruning procedure after the 3WS, the power of the 3WS approach to detect only the exact disease loci is equivalent to that of MDR with cross-validation. In the real data application, the cross-validation and 3WS analyses indicate the same two-locus model. CONCLUSIONS: Our results reveal that the performance of the two internal validation methods is equivalent with the use of pruning procedures. The specific pruning procedure should be chosen understanding the trade-off between identifying all relevant genetic effects but including false positives and missing important genetic factors. This implies 3WS may be a powerful and computationally efficient approach to screen for epistatic effects, and could be used to identify candidate interactions in large-scale genetic studies. Stacey J. Winham, Andrew J. Slater, Alison A. Motsinger-Reif |
BMC Bioinform. | 3 |
| 2008 | A balanced accuracy fitness function leads to robust analysis using grammatical evolution neural networks in the case of class imbalanceabstractGrammatical Evolution Neural Networks (GENN) is a computational method designed to detect gene-gene interactions in genetic epidemiology, but has so far only been evaluated in situations with balanced numbers of cases and controls. Real data, however, rarely has such perfectly balanced classes. In the current study, we test the power of GENN to detect interactions in data with a range of class imbalance using two fitness functions (classification error and balanced error), as well as data re-sampling. We show that when using classification error, class imbalance greatly decreases the power of GENN. Re-sampling methods demonstrated improved power, but using balanced accuracy resulted in the highest power. Based on the results of this study, balanced error has replaced classification error in the GENN algorithm. Nicholas E. Hardison, Theresa J. Fanelli, Scott M. Dudek, David M. Reif, Marylyn D. Ritchie, Alison A. Motsinger-Reif |
GECCO | 6 |
| 2007 | Linkage Disequilibrium in Genetic Association Studies Improves the Performance of Grammatical Evolution Neural NetworksabstractOne of the most important goals in genetic epidemiology is the identification of genetic factors/features that predict complex diseases. The ubiquitous nature of gene-gene interactions in the underlying etiology of common diseases creates an important analytical challenge, spurring the introduction of novel, computational approaches. One such method is a grammatical evolution neural network (GENN) approach. GENN has been shown to have high power to detect such interactions in simulation studies, but previous studies have ignored an important feature of most genetic data: linkage disequilibrium (LD). LD describes the non-random association of alleles not necessarily on the same chromosome. This results in strong correlation between variables in a dataset, which can complicate analysis. In the current study, data simulations with a range of LD patterns are used to assess the impact of such correlated variables on the performance of GENN. Our results show that not only do patterns of strong LD not decrease the power of GENN to detect genetic associations, they actually increase its power. Alison A. Motsinger-Reif, David M. Reif, Theresa J. Fanelli, Anna C. Davis, Marylyn D. Ritchie |
CIBCB | 1 |
| 2006 | Understanding the Evolutionary Process of Grammatical Evolution Neural Networks for Feature Selection in Genetic EpidemiologyabstractThe identification of genetic factors/features that predict complex diseases is an important goal of human genetics. The commonality of gene-gene interactions in the underlying genetic architecture of common diseases presents a daunting analytical challenge. Previously, we introduced a grammatical evolution neural network (GENN) approach that has high power to detect such interactions in the absence of any marginal main effects. While the success of this method is encouraging, it elicits questions regarding the evolutionary process of the algorithm itself and the feasibility of scaling the method to account for the immense dimensionality of datasets with enormous numbers of features. When the features of interest show no main effects, how is GENN able to build correct models? How and when should evolutionary parameters be adjusted according to the scale of a particular dataset? In the current study, we monitor the performance of GENN during its evolutionary process using different population sizes and numbers of generations. We also compare the evolutionary characteristics of GENN to that of a random search neural network strategy to better understand the benefits provided by the evolutionary learning process-including advantages with respect to chromosome size and the representation of functional versus non-functional features within the models generated by the two approaches. Finally, we apply lessons from the characterization of GENN to analyses of datasets containing increasing numbers of features to demonstrate the scalability of the method. Alison A. Motsinger-Reif, David M. Reif, Scott M. Dudek, Marylyn D. Ritchie |
CIBCB | 1 |
| 2006 | Feature Selection using a Random Forests Classifier for the Integrated Analysis of Multiple Data TypesabstractComplex clinical phenotypes arise from the concerted interactions among the myriad components of a biological system. Therefore, comprehensive models can only be developed through the integrated study of multiple types of experimental data gathered from the system in question. The Random Foreststrade(RF) method is adept at identifying relevant features having only slight main effects in high-dimensional data. This method is well-suited to integrated analysis, as relevant attributes may be selected from categorical or continuous data, and there may be interactions across data types. RF is a natural approach for studying gene-gene, gene-protein, or protein-protein interactions because importance scores for particular attributes take interactions into account. Thus, Random Forests is a promising solution to the analysis challenge posed by high-dimensional datasets including interactions among attributes of different types. In this study, we characterize the performance of RF on a range of simulated genetic and/or proteomic datasets. We compare the performance of RF in identifying relevant attributes when given genetic data alone, proteomic data alone, or a combined dataset of genetic plus proteomic data. Our results indicate that utilizing multiple data types is beneficial when the disease model is complex and the phenotypic outcome-associated data type is unknown. The results of this study also show that RF is adept at identifying relevant features in high-dimensional data with small main effects and low heritability David M. Reif, Alison A. Motsinger-Reif, Brett A. McKinney, James E. Crowe Jr., Jason H. Moore |
CIBCB | 2 |
| 2006 | Alternative cross-over strategies and selection techniques for grammatical evolution optimized neural networksabstractNo abstract available. Alison A. Motsinger-Reif, Lance W. Hahn, Scott M. Dudek, Kelli K. Ryckman, Marylyn D. Ritchie |
GECCO | 1 |
| 2006 | GPNN: Power studies and applications of a neural network method for detecting gene-gene interactions in studies of human diseaseabstractBACKGROUND: The identification and characterization of genes that influence the risk of common, complex multifactorial disease primarily through interactions with other genes and environmental factors remains a statistical and computational challenge in genetic epidemiology. We have previously introduced a genetic programming optimized neural network (GPNN) as a method for optimizing the architecture of a neural network to improve the identification of gene combinations associated with disease risk. The goal of this study was to evaluate the power of GPNN for identifying high-order gene-gene interactions. We were also interested in applying GPNN to a real data analysis in Parkinson's disease. RESULTS: We show that GPNN has high power to detect even relatively small genetic effects (2-3% heritability) in simulated data models involving two and three locus interactions. The limits of detection were reached under conditions with very small heritability (<1%) or when interactions involved more than three loci. We tested GPNN on a real dataset comprised of Parkinson's disease cases and controls and found a two locus interaction between the DLST gene and sex. CONCLUSION: These results indicate that GPNN may be a useful pattern recognition approach for detecting gene-gene and gene-environment interactions. Alison A. Motsinger-Reif, Stephen L. Lee, George Mellick, Marylyn D. Ritchie |
BMC Bioinform. | 1 |