Marylyn D. Ritchie

dblp:65/714 · also Marylyn DeRiggi Ritchie · DBLP profile ↗
← Back
39ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0002-1208-1720ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 31 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 A Comprehensive Benchmark of Tool-Augmented Large Language Models for Biomedical Knowledge Retrieval and Integration
abstract
General-purpose large language models (LLMs) often struggle in specialized biomedical applications due to their limited access to up-to-date, structured knowledge and domainspecific tools. Although recent studies show proof of life for AI agents in increasingly complex scientific tasks, few studies offer insights at an atomic-level. We present the first large-scale evaluation of LLM tool-calling capabilities, focused on genomics annotation tasks such as variant-to-position and variant-to-gene mapping. Our study benchmarks over one hundred LLMs, via OpenRouter's metagateway, using a standardized tool-calling protocol to retrieve information directly from biomedical APIs (e.g., NCBI dbSNP, Entrez Gene). Our experimental results show that models equipped with structured tool access significantly outperform prompt-only baselines in accuracy, factual consistency, and verifiability. These findings demonstrate the necessity of tool augmentation for reliable biomedical reasoning and provide practical insights for building and testing LLM-based agents across diverse biomedical workflows.
Van Q. Truong, Shu Yang 0009, Li Shen 0001, Marylyn D. Ritchie
BIBM4
2025 CASTER-DTA: equivariant graph neural networks for predicting drug-target affinity
abstract
Accurately determining the binding affinity of a ligand with a protein is important for drug design, development, and screening. With the advent of accessible protein structure prediction methods such as AlphaFold, predicted protein 3D structures are readily available; however, scalable methods for predicting binding affinity currently do not take full advantage of 3D protein information. Here, we present CASTER-DTA (Cross-Attention with Structural Target Equivariant Representations for Drug-Target Affinity), which uses an equivariant graph neural network (GNN) to learn more robust protein representations alongside a standard GNN to learn molecular representations to predict DTA. We augment these representations by incorporating an attention-based mechanism between protein residues and drug atoms to improve interpretability. We show that CASTER-DTA represents a state-of-the-art improvement on multiple benchmarks for predicting DTA, and that it generates novel insights for several related tasks. We then apply CASTER-DTA to create a large resource of the binding affinities of every drug approved by the U.S. Food and Drug Administration (FDA) against every protein in the human proteome and make these predictions freely available for download. We also make available a web server for researchers to apply a pretrained CASTER-DTA model for predicting binding affinities between arbitrary proteins and drugs.
Rachit Kumar, Joseph D. Romano, Marylyn D. Ritchie
Briefings Bioinform.3
2024 Genetic Algorithm Selection of Interacting Features (GASIF) for Selecting Biological Gene-Gene Interactions
abstract
Feature interactions are particularly useful in modeling biological effects, such as gene-gene interactions, but are difficult to model due to the exponential increase in the feature space. We present GASIF, a Genetic Algorithm that selects features and their interactions for the purposes of solving a supervised classification problem, designed for the identification of gene-gene interactions. GASIF works by constructing individuals with a collection of chromosomes that represent a subset of features and their interactions. It then determines individual fitness as a combination of the number of unique features used and the cross-validation performance of a logistic regression classifier trained on that feature subset with an ElasticNet penalty. A variety of intuitive operations are used to select, mate, and mutate individuals from generation to generation to limit the search space of features and interactions. We evaluate this Genetic Algorithm on a real-world dataset of human brain transcriptomic data from neuropathologically normal postmortem samples and pathologically confirmed late-onset Alzheimer's disease individuals and determine the face validity of the gene-gene interactions that it identifies. Across multiple iterations of GASIF, we consistently identified the same features and interactions as most informative, all of which relate to genes known to be implicated in Alzheimer's disease.
Rachit Kumar, Marylyn D. Ritchie
GECCO3
2024 Interpretable deep clustering survival machines for Alzheimer's disease subtype discovery
Bojian Hou, Zixuan Wen, Jingxuan Bao, Richard Zhang 0001, Boning Tong, Shu Yang 0009, Junhao Wen 0002, Yuhan Cui, Jason H. Moore, Andrew J. Saykin, Heng Huang 0001, Paul M. Thompson, Marylyn D. Ritchie, Christos Davatzikos, Li Shen 0001
Medical Image Anal.13
2022 ColocQuiaL: a QTL-GWAS colocalization pipeline
abstract
SUMMARY: Identifying genomic features responsible for genome-wide association study (GWAS) signals has proven to be a difficult challenge; many researchers have turned to colocalization analysis of GWAS signals with expression quantitative trait loci (eQTL) and splicing quantitative trait loci (sQTL) to connect GWAS signals to candidate causal genes. The ColocQuiaL pipeline provides a framework to perform these colocalization analyses at scale across the genome and returns summary files and locus visualization plots to allow for detailed review of the results. As an example, we used ColocQuiaL to perform colocalization between a recent type 2 diabetes GWAS and Genotype-Tissue Expression (GTEx) v8 single-tissue eQTL and sQTL data. AVAILABILITY AND IMPLEMENTATION: ColocQuiaL is primarily written in R and is freely available on GitHub: https://github.com/bvoightlab/ColocQuiaL.
Brian Y. Chen, William Bone, Kim Lorenz, Michael Levin 0002, Marylyn D. Ritchie, Benjamin F. Voight
Bioinform.5
2022 Estimating the effect size of a hidden causal factor between SNPs and a continuous trait: a mediation model approach
abstract
BACKGROUND: Observational studies and Mendelian randomization experiments have been used to identify many causal factors for complex traits in humans. Given a set of causal factors, it is important to understand the extent to which these causal factors explain some, all, or none of the genetic heritability, as measured by single-nucleotide polymorphisms (SNPs) that are associated with the trait. Using the mediation model framework with SNPs as the exposure, a trait of interest as the outcome, and the known causal factors as the mediators, we hypothesize that any unexplained association between the SNPs and the outcome trait is mediated by an additional unobserved, hidden causal factor. RESULTS: We propose a method to infer the effect size of this hidden mediating causal factor on the outcome trait by utilizing the estimated associations between a continuous outcome trait, the known causal factors, and the SNPs. The proposed method consists of three steps and, in the end, implements Markov chain Monte Carlo to obtain a posterior distribution for the effect size of the hidden mediator. We evaluate our proposed method via extensive simulations and show that when model assumptions hold, our method estimates the effect size of the hidden mediator well and controls type I error rate if the hidden mediator does not exist. In addition, we apply the method to the UK Biobank data and estimate parameters for a potential hidden mediator for waist-hip ratio beyond body mass index (BMI), and find that the hidden mediator has a large effect size relatively to the effect size of the known mediator BMI. CONCLUSIONS: We develop a framework to infer the effect of potential, hidden mediators influencing complex traits. This framework can be used to place boundaries on unexplained risk factors contributing to complex traits.
Zhuoran Ding, Marylyn D. Ritchie, Benjamin F. Voight, Wei-Ting Hwang
BMC Bioinform.2
2022 A research agenda to support the development and implementation of genomics-based clinical informatics tools and resources
abstract
OBJECTIVE: The Genomic Medicine Working Group of the National Advisory Council for Human Genome Research virtually hosted its 13th genomic medicine meeting titled "Developing a Clinical Genomic Informatics Research Agenda". The meeting's goal was to articulate a research strategy to develop Genomics-based Clinical Informatics Tools and Resources (GCIT) to improve the detection, treatment, and reporting of genetic disorders in clinical settings. MATERIALS AND METHODS: Experts from government agencies, the private sector, and academia in genomic medicine and clinical informatics were invited to address the meeting's goals. Invitees were also asked to complete a survey to assess important considerations needed to develop a genomic-based clinical informatics research strategy. RESULTS: Outcomes from the meeting included identifying short-term research needs, such as designing and implementing standards-based interfaces between laboratory information systems and electronic health records, as well as long-term projects, such as identifying and addressing barriers related to the establishment and implementation of genomic data exchange systems that, in turn, the research community could help address. DISCUSSION: Discussions centered on identifying gaps and barriers that impede the use of GCIT in genomic medicine. Emergent themes from the meeting included developing an implementation science framework, defining a value proposition for all stakeholders, fostering engagement with patients and partners to develop applications under patient control, promoting the use of relevant clinical workflows in research, and lowering related barriers to regulatory processes. Another key theme was recognizing pervasive biases in data and information systems, algorithms, access, value, and knowledge repositories and identifying ways to resolve them.
Ken Wiley, Laura Findley, Madison Goldrich, Teji Rakhra-Burris, Ana Stevens, Pamela Williams, Carol J. Bult, Rex L. Chisholm, Patricia Deverka, Geoffrey S. Ginsburg, Eric D. Green, Gail P. Jarvik, George A. Mensah, Erin Ramos, Mary Relling, Dan M. Roden, Robb Rowley, Gil Alterovitz, Samuel J. Aronson, Lisa Bastarache, James J. Cimino, Erin L. Crowgey, Guilherme Del Fiol, Robert R. Freimuth, Mark A. Hoffman, Janina M. Jeff, Kevin B. Johnson, Kensaku Kawamoto, Subha Madhavan, Eneida A. Mendonça, Lucila Ohno-Machado, Siddharth Pratap, Casey Overby Taylor, Marylyn D. Ritchie, Nephi Walton, Chunhua Weng, Teresa Zayas-Cabán, Teri A. Manolio, Marc S. Williams
J. Am. Medical Informatics Assoc.34
2020 Statistical Impact of Sample Size and Imbalance on Multivariate Analysis in silico and A Case Study in the UK Biobank
Xinyuan Zhang 0003, Ruowang Li, Marylyn D. Ritchie
AMIA3
2019 Real world scenarios in rare variant association analysis: the impact of imbalance and sample size on the power in silico
abstract
BACKGROUND: The development of sequencing techniques and statistical methods provides great opportunities for identifying the impact of rare genetic variation on complex traits. However, there is a lack of knowledge on the impact of sample size, case numbers, the balance of cases vs controls for both burden and dispersion based rare variant association methods. For example, Phenome-Wide Association Studies may have a wide range of case and control sample sizes across hundreds of diagnoses and traits, and with the application of statistical methods to rare variants, it is important to understand the strengths and limitations of the analyses. RESULTS: We conducted a large-scale simulation of randomly selected low-frequency protein-coding regions using twelve different balanced samples with an equal number of cases and controls as well as twenty-one unbalanced sample scenarios. We further explored statistical performance of different minor allele frequency thresholds and a range of genetic effect sizes. Our simulation results demonstrate that using an unbalanced study design has an overall higher type I error rate for both burden and dispersion tests compared with a balanced study design. Regression has an overall higher type I error with balanced cases and controls, while SKAT has higher type I error for unbalanced case-control scenarios. We also found that both type I error and power were driven by the number of cases in addition to the case to control ratio under large control group scenarios. Based on our power simulations, we observed that a SKAT analysis with case numbers larger than 200 for unbalanced case-control models yielded over 90% power with relatively well controlled type I error. To achieve similar power in regression, over 500 cases are needed. Moreover, SKAT showed higher power to detect associations in unbalanced case-control scenarios than regression. CONCLUSIONS: Our results provide important insights into rare variant association study designs by providing a landscape of type I error and statistical power for a wide range of sample sizes. These results can serve as a benchmark for making decisions about study design for rare variant analyses.
Xinyuan Zhang 0003, Anna Okula Basile, Sarah A. Pendergrass, Marylyn D. Ritchie
BMC Bioinform.4
2018 Novel features and enhancements in BioBin, a tool for the biologically inspired binning and association analysis of rare variants
abstract
Motivation: BioBin is an automated bioinformatics tool for the multi-level biological binning of sequence variants. Herein, we present a significant update to BioBin which expands the software to facilitate a comprehensive rare variant analysis and incorporates novel features and analysis enhancements. Results: In BioBin 2.3, we extend our software tool by implementing statistical association testing, updating the binning algorithm, as well as incorporating novel analysis features providing for a robust, highly customizable, and unified rare variant analysis tool. Availability and implementation: The BioBin software package is open source and freely available to users at http://www.ritchielab.com/software/biobin-download. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Anna Okula Basile, Marta Byrska-Bishop, John R. Wallace, Alex Thomas Frase, Marylyn D. Ritchie
Bioinform.5
2018 A simulation study investigating power estimates in phenome-wide association studies
abstract
BACKGROUND: Phenome-wide association studies (PheWAS) are a high-throughput approach to evaluate comprehensive associations between genetic variants and a wide range of phenotypic measures. PheWAS has varying sample sizes for quantitative traits, and variable numbers of cases and controls for binary traits across the many phenotypes of interest, which can affect the statistical power to detect associations. The motivation of this study is to investigate the various parameters which affect the estimation of statistical power in PheWAS, including sample size, case-control ratio, minor allele frequency, and disease penetrance. RESULTS: We performed a PheWAS simulation study, where we investigated variations in statistical power based on different parameters, such as overall sample size, number of cases, case-control ratio, minor allele frequency, and disease penetrance. The simulation was performed on both binary and quantitative phenotypic measures. Our simulation on binary traits suggests that the number of cases has more impact on statistical power than the case to control ratio; also, we found that a sample size of 200 cases or more maintains the statistical power to identify associations for common variants. For quantitative traits, a sample size of 1000 or more individuals performed best in the power calculations. We focused on common genetic variants (MAF > 0.01) in this study; however, in future studies, we will be extending this effort to perform similar simulations on rare variants. CONCLUSIONS: This study provides a series of PheWAS simulation analyses that can be used to estimate statistical power for some potential scenarios. These results can be used to provide guidelines for appropriate study design for future PheWAS analyses.
Anurag Verma, Yukiko Bradford, Scott M. Dudek, Anastasia Lucas, Shefali S. Verma, Sarah A. Pendergrass, Marylyn D. Ritchie
BMC Bioinform.7
2017 Using knowledge-driven genomic interactions for multi-omics data analysis: metadimensional models for predicting clinical outcomes in ovarian carcinoma
abstract
It is common that cancer patients have different molecular signatures even though they have similar clinical features, such as histology, due to the heterogeneity of tumors. To overcome this variability, we previously developed a new approach incorporating prior biological knowledge that identifies knowledge-driven genomic interactions associated with outcomes of interest. However, no systematic approach has been proposed to identify interaction models between pathways based on multi-omics data. Here we have proposed such a novel methodological framework, called metadimensional knowledge-driven genomic interactions (MKGIs). To test the utility of the proposed framework, we applied it to an ovarian cancer dataset including multi-omics profiles from The Cancer Genome Atlas to predict grade, stage, and survival outcome. We found that each knowledge-driven genomic interaction model, based on different genomic datasets, contains different sets of pathway features, which suggests that each genomic data type may contribute to outcomes in ovarian cancer via a different pathway. In addition, MKGI models significantly outperformed the single knowledge-driven genomic interaction model. From the MKGI models, many interactions between pathways associated with outcomes were found, including the mitogen-activated protein kinase (MAPK) signaling pathway and the gonadotropin-releasing hormone (GnRH) signaling pathway, which are known to play important roles in cancer pathogenesis. The beauty of incorporating biological knowledge into the model based on multi-omics data is the ability to improve diagnosis and prognosis and provide better interpretability. Thus, determining variability in molecular signatures based on these interactions between pathways may lead to better diagnostic/treatment strategies for better precision medicine.
Do Kyoon Kim, Ruowang Li, Anastasia Lucas, Shefali S. Verma, Scott M. Dudek, Marylyn D. Ritchie
J. Am. Medical Informatics Assoc.6
2016 Pathway analysis by randomization incorporating structure - PARIS: an update
abstract
MOTIVATION: We present an update to the pathway enrichment analysis tool 'Pathway Analysis by Randomization Incorporating Structure (PARIS)' that determines aggregated association signals generated from genome-wide association study results. Pathway-based analyses highlight biological pathways associated with phenotypes. PARIS uses a unique permutation strategy to evaluate the genomic structure of interrogated pathways, through permutation testing of genomic features, thus eliminating many of the over-testing concerns arising with other pathway analysis approaches. RESULTS: We have updated PARIS to incorporate expanded pathway definitions through the incorporation of new expert knowledge from multiple database sources, through customized user provided pathways, and other improvements in user flexibility and functionality. AVAILABILITY AND IMPLEMENTATION: PARIS is freely available to all users at https://ritchielab.psu.edu/software/paris-download CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Mariusz Butkiewicz, Jessica Cooke Bailey, Alex Thomas Frase, Scott M. Dudek, Brian L. Yaspan, Marylyn D. Ritchie, Sarah A. Pendergrass, Jonathan L. Haines
Bioinform.6
2015 Contrasting Association Results Between Existing PheWAS Phenotype Definition Methods and Five Validated Electronic Phenotypes
Joseph B. Leader, Sarah A. Pendergrass, Anurag Verma, David J. Carey, Dustin N. Hartzel, Marylyn D. Ritchie, H. Lester Kirchner
AMIA6
2015 Knowledge boosting: a graph-based integration approach with multi-omics data and genomic knowledge for cancer clinical outcome prediction
abstract
OBJECTIVE: Cancer can involve gene dysregulation via multiple mechanisms, so no single level of genomic data fully elucidates tumor behavior due to the presence of numerous genomic variations within or between levels in a biological system. We have previously proposed a graph-based integration approach that combines multi-omics data including copy number alteration, methylation, miRNA, and gene expression data for predicting clinical outcome in cancer. However, genomic features likely interact with other genomic features in complex signaling or regulatory networks, since cancer is caused by alterations in pathways or complete processes. METHODS: Here we propose a new graph-based framework for integrating multi-omics data and genomic knowledge to improve power in predicting clinical outcomes and elucidate interplay between different levels. To highlight the validity of our proposed framework, we used an ovarian cancer dataset from The Cancer Genome Atlas for predicting stage, grade, and survival outcomes. RESULTS: Integrating multi-omics data with genomic knowledge to construct pre-defined features resulted in higher performance in clinical outcome prediction and higher stability. For the grade outcome, the model with gene expression data produced an area under the receiver operating characteristic curve (AUC) of 0.7866. However, models of the integration with pathway, Gene Ontology, chromosomal gene set, and motif gene set consistently outperformed the model with genomic data only, attaining AUCs of 0.7873, 0.8433, 0.8254, and 0.8179, respectively. CONCLUSIONS: Integrating multi-omics data and genomic knowledge to improve understanding of molecular pathogenesis and underlying biology in cancer should improve diagnostic and prognostic indicators and the effectiveness of therapies.
Do Kyoon Kim, Je-Gun Joung, Kyung-Ah Sohn 0001, Hyunjung Shin, Yu Rang Park, Marylyn D. Ritchie, Ju Han Kim
J. Am. Medical Informatics Assoc.6
2015 Predicting censored survival data based on the interactions between meta-dimensional omics data in breast cancer
abstract
Evaluation of survival models to predict cancer patient prognosis is one of the most important areas of emphasis in cancer research. A binary classification approach has difficulty directly predicting survival due to the characteristics of censored observations and the fact that the predictive power depends on the threshold used to set two classes. In contrast, the traditional Cox regression approach has some drawbacks in the sense that it does not allow for the identification of interactions between genomic features, which could have key roles associated with cancer prognosis. In addition, data integration is regarded as one of the important issues in improving the predictive power of survival models since cancer could be caused by multiple alterations through meta-dimensional genomic data including genome, epigenome, transcriptome, and proteome. Here we have proposed a new integrative framework designed to perform these three functions simultaneously: (1) predicting censored survival data; (2) integrating meta-dimensional omics data; (3) identifying interactions within/between meta-dimensional genomic features associated with survival. In order to predict censored survival time, martingale residuals were calculated as a new continuous outcome and a new fitness function used by the grammatical evolution neural network (GENN) based on mean absolute difference of martingale residuals was implemented. To test the utility of the proposed framework, a simulation study was conducted, followed by an analysis of meta-dimensional omics data including copy number, gene expression, DNA methylation, and protein expression data in breast cancer retrieved from The Cancer Genome Atlas (TCGA). On the basis of the results from breast cancer dataset, we were able to identify interactions not only within a single dimension of genomic data but also between meta-dimensional omics data that are associated with survival. Notably, the predictive power of our best meta-dimensional model was 73% which outperformed all of the other models conducted based on a single dimension of genomic data. Breast cancer is an extremely heterogeneous disease and the high levels of genomic diversity within/between breast tumors could affect the risk of therapeutic responses and disease progression. Thus, identifying interactions within/between meta-dimensional omics data associated with survival in breast cancer is expected to deliver direction for improved meta-dimensional prognostic biomarkers and therapeutic targets.
Do Kyoon Kim, Ruowang Li, Scott M. Dudek, Marylyn D. Ritchie
J. Biomed. Informatics4
2014 Replication of SCN5A Associations with Electrocardiographic Traits in African Americans from Clinical and Epidemiologic Studies
Janina M. Jeff, Kristin Brown-Gentry, Robert J. Goodloe, Marylyn D. Ritchie, Joshua C. Denny, Abel N. Kho, Loren L. Armstrong, Bob McClellan Jr., Ping Mayo, Hailing Jin, Niloufar B. Gillani, Nathalie Schnetz-Boutaud, Holli H. Dilks, Melissa A. Basford, Jennifer A. Pacheco, Gail P. Jarvik, Rex L. Chisholm, Dan M. Roden, M. Geoffrey Hayes, Dana C. Crawford
EvoApplications4
2014 An integrated analysis of genome-wide DNA methylation and genetic variants underlying etoposide-induced cytotoxicity in European and African populations
Ruowang Li, Do Kyoon Kim, Scott M. Dudek, Marylyn D. Ritchie
EvoApplications4
2014 Benefits of Accurate Imputations in GWAS
Shefali S. Verma, Peggy L. Peissig, Deanna S. Cross, Carol Waudby, Murray H. Brilliant, Catherine A. McCarty, Marylyn D. Ritchie
EvoApplications7
2014 ATHENA: the analysis tool for heritable and environmental network associations
abstract
MOTIVATION: Advancements in high-throughput technology have allowed researchers to examine the genetic etiology of complex human traits in a robust fashion. Although genome-wide association studies have identified many novel variants associated with hundreds of traits, a large proportion of the estimated trait heritability remains unexplained. One hypothesis is that the commonly used statistical techniques and study designs are not robust to the complex etiology that may underlie these human traits. This etiology could include non-linear gene × gene or gene × environment interactions. Additionally, other levels of biological regulation may play a large role in trait variability. RESULTS: To address the need for computational tools that can explore enormous datasets to detect complex susceptibility models, we have developed a software package called the Analysis Tool for Heritable and Environmental Network Associations (ATHENA). ATHENA combines various variable filtering methods with machine learning techniques to analyze high-throughput categorical (i.e. single nucleotide polymorphisms) and quantitative (i.e. gene expression levels) predictor variables to generate multivariable models that predict either a categorical (i.e. disease status) or quantitative (i.e. cholesterol levels) outcomes. The goal of this article is to demonstrate the utility of ATHENA using simulated and biological datasets that consist of both single nucleotide polymorphisms and gene expression variables to identify complex prediction models. Importantly, this method is flexible and can be expanded to include other types of high-throughput data (i.e. RNA-seq data and biomarker measurements). AVAILABILITY: ATHENA is freely available for download. The software, user manual and tutorial can be downloaded from http://ritchielab.psu.edu/ritchielab/software.
Emily Rose Holzinger, Scott M. Dudek, Alex Thomas Frase, Sarah A. Pendergrass, Marylyn D. Ritchie
Bioinform.5
2012 A comparison of cataloged variation between International HapMap Consortium and 1000 Genomes Project data
abstract
BACKGROUND: Since publication of the human genome in 2003, geneticists have been interested in risk variant associations to resolve the etiology of traits and complex diseases. The International HapMap Consortium undertook an effort to catalog all common variation across the genome (variants with a minor allele frequency (MAF) of at least 5% in one or more ethnic groups). HapMap along with advances in genotyping technology led to genome-wide association studies which have identified common variants associated with many traits and diseases. In 2008 the 1000 Genomes Project aimed to sequence 2500 individuals and identify rare variants and 99% of variants with a MAF of <1%. METHODS: To determine whether the 1000 Genomes Project includes all the variants in HapMap, we examined the overlap between single nucleotide polymorphisms (SNPs) genotyped in the two resources using merged phase II/III HapMap data and low coverage pilot data from 1000 Genomes. RESULTS: Comparison of the two data sets showed that approximately 72% of HapMap SNPs were also found in 1000 Genomes Project pilot data. After filtering out HapMap variants with a MAF of <5% (separately for each population), 99% of HapMap SNPs were found in 1000 Genomes data. CONCLUSIONS: Not all variants cataloged in HapMap are also cataloged in 1000 Genomes. This could affect decisions about which resource to use for SNP queries, rare variant validation, or imputation. Both the HapMap and 1000 Genomes Project databases are useful resources for human genetics, but it is important to understand the assumptions made and filtering strategies employed by these projects.
Carrie C. Buchanan, Eric Torstenson, William S. Bush, Marylyn D. Ritchie
J. Am. Medical Informatics Assoc.4
2011 Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarin
abstract
OBJECTIVE: DNA biobanks linked to comprehensive electronic health records systems are potentially powerful resources for pharmacogenetic studies. This study sought to develop natural-language-processing algorithms to extract drug-dose information from clinical text, and to assess the capabilities of such tools to automate the data-extraction process for pharmacogenetic studies. MATERIALS AND METHODS: A manually validated warfarin pharmacogenetic study identified a cohort of 1125 patients with a stable warfarin dose, in which 776 patients were managed by Coumadin Clinic physicians, and the remaining 349 patients were managed by their providers. The authors developed two algorithms to extract weekly warfarin doses from both data sets: a regular expression-based program for semistructured Coumadin Clinic notes; and an advanced weekly dose calculator based on an existing medication information extraction system (MedEx) for narrative providers' notes. The authors then conducted an association analysis between an automatically extracted stable weekly dose of warfarin and four genetic variants of VKORC1 and CYP2C9 genes. The performance of the weekly dose-extraction program was evaluated by comparing it with a gold standard containing manually curated weekly doses. Precision, recall, F-measure, and overall accuracy were reported. Associations between known variants in VKORC1 and CYP2C9 and warfarin stable weekly dose were performed with linear regression adjusted for age, gender, and body mass index. RESULTS: The authors' evaluation showed that the MedEx-based system could determine patients' warfarin weekly doses with 99.7% recall, 90.8% precision, and 93.8% accuracy. Using the automatically extracted weekly doses of warfarin, the authors successfully replicated the previous known associations between warfarin stable dose and genetic variants in VKORC1 and CYP2C9.
Hua Xu 0001, Min Jiang 0007, Matthew Oetjens, Erica A. Bowton, Andrea H. Ramirez, Janina M. Jeff, Melissa A. Basford, Jill M. Pulley, James D. Cowan, Marylyn D. Ritchie, Daniel R. Masys, Dan M. Roden, Dana C. Crawford, Joshua C. Denny
J. Am. Medical Informatics Assoc.11
2010 Initialization parameter sweep in ATHENA: optimizing neural networks for detecting gene-gene interactions in the presence of small main effects
abstract
Recent advances in genotyping technology have led to the generation of an enormous quantity of genetic data. Traditional methods of statistical analysis have proved insufficient in extracting all of the information about the genetic components of common, complex human diseases. A contributing factor to the problem of analysis is that amongst the small main effects of each single gene on disease susceptibility, there are non-linear, gene-gene interactions that can be difficult for traditional, parametric analyses to detect. In addition, exhaustively searching all multi-locus combinations has proved computationally impractical. Novel strategies for analysis have been developed to address these issues. The Analysis Tool for Heritable and Environmental Network Associations (ATHENA) is an analytical tool that incorporates grammatical evolution neural networks (GENN) to detect interactions among genetic factors. Initial parameters define how the evolutionary process will be implemented. This research addresses how different parameter settings affect detection of disease models involving interactions. In the current study, we iterate over multiple parameter values to determine which combinations appear optimal for detecting interactions in simulated data for multiple genetic models. Our results indicate that the factors that have the greatest influence on detection are: input variable encoding, population size, and parallel computation.
Emily Rose Holzinger, Carrie C. Buchanan, Scott M. Dudek, Eric Torstenson, Stephen D. Turner, Marylyn D. Ritchie
GECCO6
2010 Incorporating Domain Knowledge into Evolutionary Computing for Discovering Gene-Gene Interaction
Stephen D. Turner, Scott M. Dudek, Marylyn D. Ritchie
PPSN (1)3
2010 Visualizing SNP statistics in the context of linkage disequilibrium using LD-Plus
abstract
SUMMARY: Often in human genetic analysis, multiple tables of single nucleotide polymorphism (SNP) statistics are shown alongside a Haploview style correlation plot. Readers are then asked to make inferences that incorporate knowledge across these multiple sets of results. To better facilitate a collective understanding of all available data, we developed a Ruby-based web application, LD-Plus, to generate figures that simultaneously display physical location of SNPs, binary SNP attributes (such as coding/non-coding or presence on genotyping platforms), common haplotypes and their frequencies and continuously scaled values (such as F(st), minor allele frequency, genotyping efficiency or P-values), all in the context of the D' and r(2) linkage disequilibrium structures. Combining these results into one comprehensive figure reduces dereferencing between figures and tables, and can provide unique insights into genetic features that are not clearly seen when results are partitioned across multiple figures and tables.
William S. Bush, Scott M. Dudek, Marylyn D. Ritchie
Bioinform.3
2010 PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene-disease associations
abstract
MOTIVATION: Emergence of genetic data coupled to longitudinal electronic medical records (EMRs) offers the possibility of phenome-wide association scans (PheWAS) for disease-gene associations. We propose a novel method to scan phenomic data for genetic associations using International Classification of Disease (ICD9) billing codes, which are available in most EMR systems. We have developed a code translation table to automatically define 776 different disease populations and their controls using prevalent ICD9 codes derived from EMR data. As a proof of concept of this algorithm, we genotyped the first 6005 European-Americans accrued into BioVU, Vanderbilt's DNA biobank, at five single nucleotide polymorphisms (SNPs) with previously reported disease associations: atrial fibrillation, Crohn's disease, carotid artery stenosis, coronary artery disease, multiple sclerosis, systemic lupus erythematosus and rheumatoid arthritis. The PheWAS software generated cases and control populations across all ICD9 code groups for each of these five SNPs, and disease-SNP associations were analyzed. The primary outcome of this study was replication of seven previously known SNP-disease associations for these SNPs. RESULTS: Four of seven known SNP-disease associations using the PheWAS algorithm were replicated with P-values between 2.8 x 10(-6) and 0.011. The PheWAS algorithm also identified 19 previously unknown statistical associations between these SNPs and diseases at P < 0.01. This study indicates that PheWAS analysis is a feasible method to investigate SNP-disease associations. Further evaluation is needed to determine the validity of these associations and the appropriate statistical thresholds for clinical significance. AVAILABILITY: The PheWAS software and code translation table are freely available at http://knowledgemap.mc.vanderbilt.edu/research.
Joshua C. Denny, Marylyn D. Ritchie, Melissa A. Basford, Jill M. Pulley, Lisa Bastarache, Kristin Brown-Gentry, Deede Wang, Daniel R. Masys, Dan M. Roden, Dana C. Crawford
Bioinform.2
2008 A balanced accuracy fitness function leads to robust analysis using grammatical evolution neural networks in the case of class imbalance
abstract
Grammatical Evolution Neural Networks (GENN) is a computational method designed to detect gene-gene interactions in genetic epidemiology, but has so far only been evaluated in situations with balanced numbers of cases and controls. Real data, however, rarely has such perfectly balanced classes. In the current study, we test the power of GENN to detect interactions in data with a range of class imbalance using two fitness functions (classification error and balanced error), as well as data re-sampling. We show that when using classification error, class imbalance greatly decreases the power of GENN. Re-sampling methods demonstrated improved power, but using balanced accuracy resulted in the highest power. Based on the results of this study, balanced error has replaced classification error in the GENN algorithm.
Nicholas E. Hardison, Theresa J. Fanelli, Scott M. Dudek, David M. Reif, Marylyn D. Ritchie, Alison A. Motsinger-Reif
GECCO5
2008 Alternative contingency table measures improve the power and detection of multifactor dimensionality reduction
abstract
BACKGROUND: Multifactor Dimensionality Reduction (MDR) has been introduced previously as a non-parametric statistical method for detecting gene-gene interactions. MDR performs a dimensional reduction by assigning multi-locus genotypes to either high- or low-risk groups and measuring the percentage of cases and controls incorrectly labelled by this classification - the classification error. The combination of variables that produces the lowest classification error is selected as the best or most fit model. The correctly and incorrectly labelled cases and controls can be expressed as a two-way contingency table. We sought to improve the ability of MDR to detect gene-gene interactions by replacing classification error with a different measure to score model quality. RESULTS: In this study, we compare the detection and power of MDR using a variety of measures for two-way contingency table analysis. We simulated 40 genetic models, varying the number of disease loci in the model (2 - 5), allele frequencies of the disease loci (.2/.8 or .4/.6) and the broad-sense heritability of the model (.05 - .3). Overall, detection using NMI was 65.36% across all models, and specific detection was 59.4% versus detection using classification error at 62% and specific detection was 52.2%. CONCLUSION: Of the 10 measures evaluated, the likelihood ratio and normalized mutual information (NMI) are measures that consistently improve the detection and power of MDR in simulated data over using classification error. These measures also reduce the inclusion of spurious variables in a multi-locus model. Thus, MDR, which has already been demonstrated as a powerful tool for detecting gene-gene interactions, can be improved with the use of alternative fitness functions.
William S. Bush, Todd L. Edwards, Scott M. Dudek, Brett A. McKinney, Marylyn D. Ritchie
BMC Bioinform.5
2007 Linkage Disequilibrium in Genetic Association Studies Improves the Performance of Grammatical Evolution Neural Networks
abstract
One of the most important goals in genetic epidemiology is the identification of genetic factors/features that predict complex diseases. The ubiquitous nature of gene-gene interactions in the underlying etiology of common diseases creates an important analytical challenge, spurring the introduction of novel, computational approaches. One such method is a grammatical evolution neural network (GENN) approach. GENN has been shown to have high power to detect such interactions in simulation studies, but previous studies have ignored an important feature of most genetic data: linkage disequilibrium (LD). LD describes the non-random association of alleles not necessarily on the same chromosome. This results in strong correlation between variables in a dataset, which can complicate analysis. In the current study, data simulations with a range of LD patterns are used to assess the impact of such correlated variables on the performance of GENN. Our results show that not only do patterns of strong LD not decrease the power of GENN to detect genetic associations, they actually increase its power.
Alison A. Motsinger-Reif, David M. Reif, Theresa J. Fanelli, Anna C. Davis, Marylyn D. Ritchie
CIBCB5
2007 Association Rule Discovery Has the Ability to Model Complex Genetic Effects
abstract
Dramatic advances in genotyping technology have established a need for fast, flexible analysis methods for genetic association studies. Common complex diseases, such as Parkinson's disease or multiple sclerosis, are thought to involve an interplay of multiple genes working either independently or together to influence disease risk. Also, multiple underlying traits, each its own genetic basis may be defined together as a single disease. These effects - trait heterogeneity, locus heterogeneity, and gene-gene interactions (epistasis) - contribute to the complex architecture of common genetic diseases. Association Rule Discovery (ARD) searches for frequent itemsets to identify rule-based patterns in large scale data. In this study, we apply Apriori (an ARD algorithm) to simulated genetic data with varying degrees of complexity. Apriori using information difference to prior as a rule measure shows good power to detect functional effects in simulated cases of simple trait heterogeneity, trait heterogeneity and epistasis, and moderate power in cases of trait heterogeneity and locus heterogeneity. Also, we illustrate that bootstrapping the rule induction process does not considerably improve the power to detect these effects. These results show that ARD is a framework with sufficient flexibility to characterize complex genetic effects.
William S. Bush, Tricia A. Thornton-Wells, Marylyn D. Ritchie
CIDM3
2006 Understanding the Evolutionary Process of Grammatical Evolution Neural Networks for Feature Selection in Genetic Epidemiology
abstract
The identification of genetic factors/features that predict complex diseases is an important goal of human genetics. The commonality of gene-gene interactions in the underlying genetic architecture of common diseases presents a daunting analytical challenge. Previously, we introduced a grammatical evolution neural network (GENN) approach that has high power to detect such interactions in the absence of any marginal main effects. While the success of this method is encouraging, it elicits questions regarding the evolutionary process of the algorithm itself and the feasibility of scaling the method to account for the immense dimensionality of datasets with enormous numbers of features. When the features of interest show no main effects, how is GENN able to build correct models? How and when should evolutionary parameters be adjusted according to the scale of a particular dataset? In the current study, we monitor the performance of GENN during its evolutionary process using different population sizes and numbers of generations. We also compare the evolutionary characteristics of GENN to that of a random search neural network strategy to better understand the benefits provided by the evolutionary learning process-including advantages with respect to chromosome size and the representation of functional versus non-functional features within the models generated by the two approaches. Finally, we apply lessons from the characterization of GENN to analyses of datasets containing increasing numbers of features to demonstrate the scalability of the method.
Alison A. Motsinger-Reif, David M. Reif, Scott M. Dudek, Marylyn D. Ritchie
CIBCB4
2006 Alternative cross-over strategies and selection techniques for grammatical evolution optimized neural networks
abstract
No abstract available.
Alison A. Motsinger-Reif, Lance W. Hahn, Scott M. Dudek, Kelli K. Ryckman, Marylyn D. Ritchie
GECCO5
2006 Parallel multifactor dimensionality reduction: a tool for the large-scale analysis of gene-gene interactions
abstract
UNLABELLED: Parallel multifactor dimensionality reduction is a tool for large-scale analysis of gene-gene and gene-environment interactions. The MDR algorithm was redesigned to allow an unlimited number of study subjects, total variables and variable states, and to remove restrictions on the order of interactions being analyzed. In addition, the algorithm is markedly more efficient, with approximately 150-fold decrease in runtime for equivalent analyses. To facilitate the processing of large datasets, the algorithm was made parallel. AVAILABILITY: Parallel MDR is freely available for non-commercial research institutions. For full details see http://chgr.mc.vanderbilt.edu/ritchielab/pMDR. An open-source version of MDR software is available at http://www.epistasis.org.
William S. Bush, Scott M. Dudek, Marylyn D. Ritchie
Bioinform.3
2006 GPNN: Power studies and applications of a neural network method for detecting gene-gene interactions in studies of human disease
abstract
BACKGROUND: The identification and characterization of genes that influence the risk of common, complex multifactorial disease primarily through interactions with other genes and environmental factors remains a statistical and computational challenge in genetic epidemiology. We have previously introduced a genetic programming optimized neural network (GPNN) as a method for optimizing the architecture of a neural network to improve the identification of gene combinations associated with disease risk. The goal of this study was to evaluate the power of GPNN for identifying high-order gene-gene interactions. We were also interested in applying GPNN to a real data analysis in Parkinson's disease. RESULTS: We show that GPNN has high power to detect even relatively small genetic effects (2-3% heritability) in simulated data models involving two and three locus interactions. The limits of detection were reached under conditions with very small heritability (<1%) or when interactions involved more than three loci. We tested GPNN on a real dataset comprised of Parkinson's disease cases and controls and found a two locus interaction between the DLST gene and sex. CONCLUSION: These results indicate that GPNN may be a useful pattern recognition approach for detecting gene-gene and gene-environment interactions.
Alison A. Motsinger-Reif, Stephen L. Lee, George Mellick, Marylyn D. Ritchie
BMC Bioinform.4
2004 Genetic Programming Neural Networks as a Bioinformatics Tool for Human Genetics
Marylyn D. Ritchie, Christopher S. Coffey, Jason H. Moore
GECCO (1)1
2004 An application of conditional logistic regression and multifactor dimensionality reduction for detecting gene-gene Interactions on risk of myocardial infarction: The importance of model validation
abstract
BACKGROUND: To examine interactions among the angiotensin converting enzyme (ACE) insertion/deletion, plasminogen activator inhibitor-1 (PAI-1) 4G/5G, and tissue plasminogen activator (t-PA) insertion/deletion gene polymorphisms on risk of myocardial infarction using data from 343 matched case-control pairs from the Physicians Health Study. We examined the data using both conditional logistic regression and the multifactor dimensionality reduction (MDR) method. One advantage of the MDR method is that it provides an internal prediction error for validation. We summarize our use of this internal prediction error for model validation. RESULTS: The overall results for the two methods were consistent, with both suggesting an interaction between the ACE I/D and PAI-1 4G/5G polymorphisms. However, using ten-fold cross validation, the 46% prediction error for the final MDR model was not significantly lower than that expected by chance. CONCLUSIONS: The significant interaction initially observed does not validate and may represent a type I error. As data-driven analytic methods continue to be developed and used to examine complex genetic interactions, it will become increasingly important to stress model validation in order to ensure that significant effects represent true relationships rather than chance findings.
Christopher S. Coffey, Patricia R. Hebert, Marylyn D. Ritchie, Harlan M. Krumholz, John Michael Gaziano, Paul M. Ridker, Nancy J. Brown, Douglas E. Vaughan, Jason H. Moore
BMC Bioinform.3
2003 Multifactor dimensionality reduction software for detecting gene-gene and gene-environment interactions
abstract
MOTIVATION: Polymorphisms in human genes are being described in remarkable numbers. Determining which polymorphisms and which environmental factors are associated with common, complex diseases has become a daunting task. This is partly because the effect of any single genetic variation will likely be dependent on other genetic variations (gene-gene interaction or epistasis) and environmental factors (gene-environment interaction). Detecting and characterizing interactions among multiple factors is both a statistical and a computational challenge. To address this problem, we have developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe the MDR approach and an MDR software package. RESULTS: We developed a program that integrates MDR with a cross-validation strategy for estimating the classification and prediction error of multifactor models. The software can be used to analyze interactions among 2-15 genetic and/or environmental factors. The dataset may contain up to 500 total variables and a maximum of 4000 study subjects. AVAILABILITY: Information on obtaining the executable code, example data, example analysis, and documentation is available upon request. SUPPLEMENTARY INFORMATION: All supplementary information can be found at http://phg.mc.vanderbilt.edu/Software/MDR.
Lance W. Hahn, Marylyn D. Ritchie, Jason H. Moore
Bioinform.2
2003 Optimizationof neural network architecture using genetic programming improvesdetection and modeling of gene-gene interactions in studies of humandiseases
abstract
BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.
Marylyn D. Ritchie, Bill C. White, Joel S. Parker, Lance W. Hahn, Jason H. Moore
BMC Bioinform.1
2002 Application Of Genetic Algorithms To The Discovery Of Complex Models For Simulation Studies In Human Genetics
Jason H. Moore, Lance W. Hahn, Marylyn D. Ritchie, Tricia A. Thornton-Wells, Bill C. White
GECCO3