EDBT 2026 Demo / reviewers in the wild / expert
Casey S. Greene
dblp:69/5821
· DBLP profile ↗
24ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0001-8713-9213ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BuDDI: Bulk Deconvolution with Domain Invariance to predict cell-type-specific perturbations from bulkabstractWhile single-cell experiments provide deep cellular resolution within a single sample, some single-cell experiments are inherently more challenging than bulk experiments due to dissociation difficulties, cost, or limited tissue availability. This creates a situation where we have deep cellular profiles of one sample or condition, and bulk profiles across multiple samples and conditions. To bridge this gap, we propose BuDDI (BUlk Deconvolution with Domain Invariance). BuDDI utilizes domain adaptation techniques to effectively integrate available corpora of case-control bulk and reference scRNA-seq observations to infer cell-type-specific perturbation effects. BuDDI achieves this by learning independent latent spaces within a single variational autoencoder (VAE) encompassing at least four sources of variability: 1) cell type proportion, 2) perturbation effect, 3) structured experimental variability, and 4) remaining variability. Since each latent space is encouraged to be independent, we simulate perturbation responses by independently composing each latent space to simulate cell-type-specific perturbation responses. We evaluated BuDDI's performance on simulated and real data with experimental designs of increasing complexity. We first validated that BuDDI could learn domain invariant latent spaces on data with matched samples across each source of variability. Then we validated that BuDDI could accurately predict cell-type-specific perturbation response when no single-cell perturbed profiles were used during training; instead, only bulk samples had both perturbed and non-perturbed observations. Finally, we validated BuDDI on predicting sex-specific differences, an experimental design where it is not possible to have matched samples. In each experiment, BuDDI outperformed all other comparative methods and baselines. As more reference atlases are completed, BuDDI provides a path to combine these resources with bulk-profiled treatment or disease signatures to study perturbations, sex differences, or other factors at single-cell resolution. Natalie R. Davidson, Fan Zhang 0133, Casey S. Greene |
PLoS Comput. Biol. | 3 |
| 2024 | A publishing infrastructure for Artificial Intelligence (AI)-assisted academic authoringabstractOBJECTIVE: Investigate the use of advanced natural language processing models to streamline the time-consuming process of writing and revising scholarly manuscripts. MATERIALS AND METHODS: For this purpose, we integrate large language models into the Manubot publishing ecosystem to suggest revisions for scholarly texts. Our AI-based revision workflow employs a prompt generator that incorporates manuscript metadata into templates, generating section-specific instructions for the language model. The model then generates revised versions of each paragraph for human authors to review. We evaluated this methodology through 5 case studies of existing manuscripts, including the revision of this manuscript. RESULTS: Our results indicate that these models, despite some limitations, can grasp complex academic concepts and enhance text quality. All changes to the manuscript are tracked using a version control system, ensuring transparency in distinguishing between human- and machine-generated text. CONCLUSIONS: Given the significant time researchers invest in crafting prose, incorporating large language models into the scholarly writing process can significantly improve the type of knowledge work performed by academics. Our approach also enables scholars to concentrate on critical aspects of their work, such as the novelty of their ideas, while automating tedious tasks like adhering to specific writing styles. Although the use of AI-assisted tools in scientific authoring is controversial, our approach, which focuses on revising human-written text and provides change-tracking transparency, can mitigate concerns regarding AI's role in scientific writing. Milton Pividori, Casey S. Greene |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | The effect of non-linear signal in classification problems using gene expressionabstractThose building predictive models from transcriptomic data are faced with two conflicting perspectives. The first, based on the inherent high dimensionality of biological systems, supposes that complex non-linear models such as neural networks will better match complex biological systems. The second, imagining that complex systems will still be well predicted by simple dividing lines prefers linear models that are easier to interpret. We compare multi-layer neural networks and logistic regression across multiple prediction tasks on GTEx and Recount3 datasets and find evidence in favor of both possibilities. We verified the presence of non-linear signal when predicting tissue and metadata sex labels from expression data by removing the predictive linear signal with Limma, and showed the removal ablated the performance of linear methods but not non-linear ones. However, we also found that the presence of non-linear signal was not necessarily sufficient for neural networks to outperform logistic regression. Our results demonstrate that while multi-layer neural networks may be useful for making predictions from gene expression data, including a linear baseline model is critical because while biological systems are high-dimensional, effective dividing lines for predictive models may not be. Benjamin J. Heil, Jake Crawford, Casey S. Greene |
PLoS Comput. Biol. | 3 |
| 2022 | wenda_gpu: fast domain adaptation for genomic dataabstractMOTIVATION: Domain adaptation allows for the development of predictive models even in cases with limited sample data. Weighted elastic net domain adaptation specifically leverages features of genomic data to maximize transferability but the method is too computationally demanding to apply to many genome-sized datasets. RESULTS: We developed wenda_gpu, which uses GPyTorch to train models on genomic data within hours on a single GPU-enabled machine. We show that wenda_gpu returns comparable results to the original wenda implementation, and that it can be used for improved prediction of cancer mutation status on small sample sizes than regular elastic net. AVAILABILITY AND IMPLEMENTATION: wenda_gpu is available on GitHub at https://github.com/greenelab/wenda_gpu/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ariel A. Hippen, Jake Crawford, Jacob R. Gardner, Casey S. Greene |
Bioinform. | 4 |
| 2022 | Ten simple rules for large-scale data processingabstractExabytes of images, sequences, tabular data, and unstructured data are now available for analysis to advance science.These data support efforts to visualize the image of the black hole [1], characterize the spectrum of mutations across cancer types [2], create a sophisticated language model [3], and many other tasks.Defining a data analysis as large scale or not is likely to be a moving target as computing and data transfer technologies advance.We consider an analysis large scale when the data size exceeds the capacity of local resources or when the amount of time spent waiting for high-performance computing (HPC) compute capacity would disrupt the pace of the research project.For example, the recount2 [4] analysis processed petabytes of data, so we consider it to be large-scale data processing for the current day standard.We provide 10 simple rules that apply whether the infrastructure used is traditional HPC or a more modern cloud service.Our work and experience are in the space of genomics, but the 10 rules we provide here are more general and broadly applicable given our definition of largescale data analysis.The particular characteristics of genomics data that influence our recommendations but that may not be universally true are that the full set of such data does not have to be present to begin processing, that many processing steps can be completed in a streaming manner, and that the size of processed data required for downstream analysis and interpretation is typically orders of magnitude smaller than the unprocessed form. Arkarachai Fungtammasan, Alexandra Lee, Jaclyn N. Taroni, Kurt Wheeler, Chen-Shan Chin, Sean R. Davis, Casey S. Greene |
PLoS Comput. Biol. | 7 |
| 2022 | Ten quick tips for deep learning in biologyabstractMachine learning is a modern approach to problem-solving and task automation.In particular, machine learning is concerned with the development and applications of algorithms that Benjamin D. Lee, Anthony Gitter, Casey S. Greene, Sebastian Raschka, Finlay Maguire, Alexander J. Titus, Michael D. Kessler, Alexandra Lee, Marc G. Chevrette, Paul Allen Stewart, Thiago Britto-Borges, Evan M. Cofer, Kun-Hsing Yu, Juan Jose Carmona, Elana J. Fertig, Alexandr A. Kalinin, Brandon Signal, Benjamin J. Lengerich, Timothy J. Triche Jr., Simina M. Boca |
PLoS Comput. Biol. | 3 |
| 2021 | miQC: An adaptive probabilistic framework for quality control of single-cell RNA-sequencing dataabstractSingle-cell RNA-sequencing (scRNA-seq) has made it possible to profile gene expression in tissues at high resolution. An important preprocessing step prior to performing downstream analyses is to identify and remove cells with poor or degraded sample quality using quality control (QC) metrics. Two widely used QC metrics to identify a 'low-quality' cell are (i) if the cell includes a high proportion of reads that map to mitochondrial DNA (mtDNA) encoded genes and (ii) if a small number of genes are detected. Current best practices use these QC metrics independently with either arbitrary, uniform thresholds (e.g. 5%) or biological context-dependent (e.g. species) thresholds, and fail to jointly model these metrics in a data-driven manner. Current practices are often overly stringent and especially untenable on certain types of tissues, such as archived tumor tissues, or tissues associated with mitochondrial function, such as kidney tissue [1]. We propose a data-driven QC metric (miQC) that jointly models both the proportion of reads mapping to mtDNA genes and the number of detected genes with mixture models in a probabilistic framework to predict the low-quality cells in a given dataset. We demonstrate how our QC metric easily adapts to different types of single-cell datasets to remove low-quality cells while preserving high-quality cells that can be used for downstream analyses. Our software package is available at https://bioconductor.org/packages/miQC. Ariel A. Hippen, Matias M. Falco, Lukas M. Weber, Erdogan Pekcan Erkan, Kaiyang Zhang, Jennifer A. Doherty, Anna Vähärautio, Casey S. Greene, Stephanie C. Hicks |
PLoS Comput. Biol. | 8 |
| 2019 | Learning and Imputation for Mass-spec Bias Reduction (LIMBR)abstractMOTIVATION: Decreasing costs are making it feasible to perform time series proteomics and genomics experiments with more replicates and higher resolution than ever before. With more replicates and time points, proteome and genome-wide patterns of expression are more readily discernible. These larger experiments require more batches exacerbating batch effects and increasing the number of bias trends. In the case of proteomics, where methods frequently result in missing data this increasing scale is also decreasing the number of peptides observed in all samples. The sources of batch effects and missing data are incompletely understood necessitating novel techniques. RESULTS: Here we show that by exploiting the structure of time series experiments, it is possible to accurately and reproducibly model and remove batch effects. We implement Learning and Imputation for Mass-spec Bias Reduction (LIMBR) software, which builds on previous block-based models of batch effects and includes features specific to time series and circadian studies. To aid in the analysis of time series proteomics experiments, which are often plagued with missing data points, we also integrate an imputation system. By building LIMBR for imputation and time series tailored bias modeling into one straightforward software package, we expect that the quality and ease of large-scale proteomics and genomics time series experiments will be significantly increased. AVAILABILITY AND IMPLEMENTATION: Python code and documentation is available for download at https://github.com/aleccrowell/LIMBR and LIMBR can be downloaded and installed with dependencies using 'pip install limbr'. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Alexander M. Crowell, Casey S. Greene, Jennifer J. Loros, Jay C. Dunlap |
Bioinform. | 2 |
| 2019 | Open collaborative writing with ManubotabstractOpen, collaborative research is a powerful paradigm that can immensely strengthen the scientific process by integrating broad and diverse expertise. However, traditional research and multi-author writing processes break down at scale. We present new software named Manubot, available at https://manubot.org, to address the challenges of open scholarly writing. Manubot adopts the contribution workflow used by many large-scale open source software projects to enable collaborative authoring of scholarly manuscripts. With Manubot, manuscripts are written in Markdown and stored in a Git repository to precisely track changes over time. By hosting manuscript repositories publicly, such as on GitHub, multiple authors can simultaneously propose and review changes. A cloud service automatically evaluates proposed changes to catch errors. Publication with Manubot is continuous: When a manuscript's source changes, the rendered outputs are rebuilt and republished to a web page. Manubot automates bibliographic tasks by implementing citation by identifier, where users cite persistent identifiers (e.g. DOIs, PubMed IDs, ISBNs, URLs), whose metadata is then retrieved and converted to a user-specified style. Manubot modernizes publishing to align with the ideals of open science by making it transparent, reproducible, immediate, versioned, collaborative, and free of charge. Daniel S. Himmelstein, Vincent Rubinetti, David R. Slochower, Dongbo Hu, Venkat S. Malladi, Casey S. Greene, Anthony Gitter |
PLoS Comput. Biol. | 6 |
| 2017 | Tissue-specific network-based genome wide study of amygdala imaging phenotypes to identify functional interaction modulesabstractMOTIVATION: Network-based genome-wide association studies (GWAS) aim to identify functional modules from biological networks that are enriched by top GWAS findings. Although gene functions are relevant to tissue context, most existing methods analyze tissue-free networks without reflecting phenotypic specificity. RESULTS: We propose a novel module identification framework for imaging genetic studies using the tissue-specific functional interaction network. Our method includes three steps: (i) re-prioritize imaging GWAS findings by applying machine learning methods to incorporate network topological information and enhance the connectivity among top genes; (ii) detect densely connected modules based on interactions among top re-prioritized genes; and (iii) identify phenotype-relevant modules enriched by top GWAS findings. We demonstrate our method on the GWAS of [18F]FDG-PET measures in the amygdala region using the imaging genetic data from the Alzheimer's Disease Neuroimaging Initiative, and map the GWAS results onto the amygdala-specific functional interaction network. The proposed network-based GWAS method can effectively detect densely connected modules enriched by top GWAS findings. Tissue-specific functional network can provide precise context to help explore the collective effects of genes with biologically meaningful interactions specific to the studied phenotype. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at http://www.iu.edu/shenlab/tools/gwasmodule/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaohui Yao, Kefei Liu 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Casey S. Greene, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 7 |
| 2017 | ADAGE signature analysis: differential expression analysis with data-defined gene setsabstractBACKGROUND: Gene set enrichment analysis and overrepresentation analyses are commonly used methods to determine the biological processes affected by a differential expression experiment. This approach requires biologically relevant gene sets, which are currently curated manually, limiting their availability and accuracy in many organisms without extensively curated resources. New feature learning approaches can now be paired with existing data collections to directly extract functional gene sets from big data. RESULTS: Here we introduce a method to identify perturbed processes. In contrast with methods that use curated gene sets, this approach uses signatures extracted from public expression data. We first extract expression signatures from public data using ADAGE, a neural network-based feature extraction approach. We next identify signatures that are differentially active under a given treatment. Our results demonstrate that these signatures represent biological processes that are perturbed by the experiment. Because these signatures are directly learned from data without supervision, they can identify uncurated or novel biological processes. We implemented ADAGE signature analysis for the bacterial pathogen Pseudomonas aeruginosa. For the convenience of different user groups, we implemented both an R package (ADAGEpath) and a web server ( http://adage.greenelab.com ) to run these analyses. Both are open-source to allow easy expansion to other organisms or signature generation methods. We applied ADAGE signature analysis to an example dataset in which wild-type and ∆anr mutant cells were grown as biofilms on the Cystic Fibrosis genotype bronchial epithelial cells. We mapped active signatures in the dataset to KEGG pathways and compared with pathways identified using GSEA. The two approaches generally return consistent results; however, ADAGE signature analysis also identified a signature that revealed the molecularly supported link between the MexT regulon and Anr. CONCLUSIONS: We designed ADAGE signature analysis to perform gene set analysis using data-defined functional gene signatures. This approach addresses an important gap for biologists studying non-traditional model organisms and those without extensive curated resources available. We built both an R package and web server to provide ADAGE signature analysis to the community. Matt Huyck, Dongbo Hu, René A. Zelaya, Deborah A. Hogan, Casey S. Greene |
BMC Bioinform. | 6 |
| 2016 | Recent Advances and Emerging Applications in Text and Data Mining for Biomedical DiscoveryabstractPrecision medicine will revolutionize the way we treat and prevent disease. A major barrier to the implementation of precision medicine that clinicians and translational scientists face is understanding the underlying mechanisms of disease. We are starting to address this challenge through automatic approaches for information extraction, representation and analysis. Recent advances in text and data mining have been applied to a broad spectrum of key biomedical questions in genomics, pharmacogenomics and other fields. We present an overview of the fundamental methods for text and data mining, as well as recent advances and emerging applications toward precision medicine. Graciela Gonzalez-Hernandez, Tasnia Tahsin, Britton C. Goodale, Anna C. Greene, Casey S. Greene |
Briefings Bioinform. | 5 |
| 2016 | Adapting bioinformatics curricula for big dataabstractModern technologies are capable of generating enormous amounts of data that measure complex biological systems. Computational biologists and bioinformatics scientists are increasingly being asked to use these data to reveal key systems-level properties. We review the extent to which curricula are changing in the era of big data. We identify key competencies that scientists dealing with big data are expected to possess across fields, and we use this information to propose courses to meet these growing needs. While bioinformatics programs have traditionally trained students in data-intensive science, we identify areas of particular biological, computational and statistical emphasis important for this era that can be incorporated into existing curricula. For each area, we propose a course structured around these topics, which can be adapted in whole or in parts into existing curricula. In summary, specific challenges associated with big data provide an important opportunity to update existing curricula, but we do not foresee a wholesale redesign of bioinformatics training programs. Anna C. Greene, Kristine A. Giffin, Casey S. Greene, Jason H. Moore |
Briefings Bioinform. | 3 |
| 2016 | Semi-supervised learning of the electronic health record for phenotype stratificationabstractPatient interactions with health care providers result in entries to electronic health records (EHRs). EHRs were built for clinical and billing purposes but contain many data points about an individual. Mining these records provides opportunities to extract electronic phenotypes, which can be paired with genetic data to identify genes underlying common human diseases. This task remains challenging: high quality phenotyping is costly and requires physician review; many fields in the records are sparsely filled; and our definitions of diseases are continuing to improve over time. Here we develop and evaluate a semi-supervised learning method for EHR phenotype extraction using denoising autoencoders for phenotype stratification. By combining denoising autoencoders with random forests we find classification improvements across multiple simulation models and improved survival prediction in ALS clinical trial data. This is particularly evident in cases where only a small number of patients have high quality phenotypes, a common scenario in EHR-based research. Denoising autoencoders perform dimensionality reduction enabling visualization and clustering for the discovery of new subtypes of disease. This method represents a promising approach to clarify disease subtypes and improve genotype-phenotype association studies that leverage EHRs. Brett K. Beaulieu-Jones, Casey S. Greene |
J. Biomed. Informatics | 2 |
| 2015 | Systems Level Analysis of Systemic Sclerosis Shows a Network of Immune and Profibrotic Pathways Connected with Genetic PolymorphismsabstractSystemic sclerosis (SSc) is a rare systemic autoimmune disease characterized by skin and organ fibrosis. The pathogenesis of SSc and its progression are poorly understood. The SSc intrinsic gene expression subsets (inflammatory, fibroproliferative, normal-like, and limited) are observed in multiple clinical cohorts of patients with SSc. Analysis of longitudinal skin biopsies suggests that a patient's subset assignment is stable over 6-12 months. Genetically, SSc is multi-factorial with many genetic risk loci for SSc generally and for specific clinical manifestations. Here we identify the genes consistently associated with the intrinsic subsets across three independent cohorts, show the relationship between these genes using a gene-gene interaction network, and place the genetic risk loci in the context of the intrinsic subsets. To identify gene expression modules common to three independent datasets from three different clinical centers, we developed a consensus clustering procedure based on mutual information of partitions, an information theory concept, and performed a meta-analysis of these genome-wide gene expression datasets. We created a gene-gene interaction network of the conserved molecular features across the intrinsic subsets and analyzed their connections with SSc-associated genetic polymorphisms. The network is composed of distinct, but interconnected, components related to interferon activation, M2 macrophages, adaptive immunity, extracellular matrix remodeling, and cell proliferation. The network shows extensive connections between the inflammatory- and fibroproliferative-specific genes. The network also shows connections between these subset-specific genes and 30 SSc-associated polymorphic genes including STAT4, BLK, IRF7, NOTCH4, PLAUR, CSK, IRAK1, and several human leukocyte antigen (HLA) genes. Our analyses suggest that the gene expression changes underlying the SSc subsets may be long-lived, but mechanistically interconnected and related to a patients underlying genetic risk. J. Matthew Mahoney, Jaclyn N. Taroni, Viktor Martyanov, Tammara A. Wood, Casey S. Greene, Patricia A. Pioli, Monique Hinchcliff, Michael L. Whitfield |
PLoS Comput. Biol. | 5 |
| 2013 | Functional Knowledge Transfer for High-accuracy Prediction of Under-studied Biological ProcessesabstractA key challenge in genetics is identifying the functional roles of genes in pathways. Numerous functional genomics techniques (e.g. machine learning) that predict protein function have been developed to address this question. These methods generally build from existing annotations of genes to pathways and thus are often unable to identify additional genes participating in processes that are not already well studied. Many of these processes are well studied in some organism, but not necessarily in an investigator's organism of interest. Sequence-based search methods (e.g. BLAST) have been used to transfer such annotation information between organisms. We demonstrate that functional genomics can complement traditional sequence similarity to improve the transfer of gene annotations between organisms. Our method transfers annotations only when functionally appropriate as determined by genomic data and can be used with any prediction algorithm to combine transferred gene function knowledge with organism-specific high-throughput data to enable accurate function prediction. We show that diverse state-of-art machine learning algorithms leveraging functional knowledge transfer (FKT) dramatically improve their accuracy in predicting gene-pathway membership, particularly for processes with little experimental knowledge in an organism. We also show that our method compares favorably to annotation transfer by sequence similarity. Next, we deploy FKT with state-of-the-art SVM classifier to predict novel genes to 11,000 biological processes across six diverse organisms and expand the coverage of accurate function predictions to processes that are often ignored because of a dearth of annotated genes in an organism. Finally, we perform in vivo experimental investigation in Danio rerio and confirm the regulatory role of our top predicted novel gene, wnt5b, in leftward cell migration during heart development. FKT is immediately applicable to many bioinformatics techniques and will help biologists systematically integrate prior knowledge from diverse systems to direct targeted experiments in their organism of study. Christopher Y. Park, Aaron K. Wong, Casey S. Greene, Jessica Rowland, Yuanfang Guan, Lars Ailo Bongo, Rebecca D. Burdine, Olga G. Troyanskaya |
PLoS Comput. Biol. | 3 |
| 2012 | Chapter 2: Data-Driven View of Disease BiologyabstractModern experimental strategies often generate genome-scale measurements of human tissues or cell lines in various physiological states. Investigators often use these datasets individually to help elucidate molecular mechanisms of human diseases. Here we discuss approaches that effectively weight and integrate hundreds of heterogeneous datasets to gene-gene networks that focus on a specific process or disease. Diverse and systematic genome-scale measurements provide such approaches both a great deal of power and a number of challenges. We discuss some such challenges as well as methods to address them. We also raise important considerations for the assessment and evaluation of such approaches. When carefully applied, these integrative data-driven methods can make novel high-quality predictions that can transform our understanding of the molecular-basis of human disease. Casey S. Greene, Olga G. Troyanskaya |
PLoS Comput. Biol. | 1 |
| 2010 | Fast genome-wide epistasis analysis using ant colony optimization for multifactor dimensionality reduction analysis on graphics processing unitsabstractEpistasis, or non-linear gene-to-gene interaction, is now thought to be at the heart of many common human diseases. A popular algorithm to detect epistasis is Multifactor Dimensionality Reduction (MDR), which exhaustively searches to determine an optimal classification. This exhaustive search is combinatorial in complexity and does not scale efficiently to large datasets. Ant Colony Opimization (ACO) is a technique to reduce this complexity by exploiting expert knowledge to spend more time looking at most likely candidates for the optimal classification. Graphics Processing Units (GPUs) are highly-parallel integrated circuits able to execute arbitrary code. The authors implemented ACO MDR on GPUs and compared it to both a Java ACO implementation and an exhaustive C++ implementation. The performance advantage of GPUs, combined with the added computational efficiency of a heuristic evolutionary algorithm such as ACO, allow larger scale problems to be tackled, something that is becoming critical with the advances in high throughput genome sequencing. Nicholas A. Sinnott-Armstrong, Casey S. Greene, Jason H. Moore |
GECCO | 2 |
| 2010 | Multifactor dimensionality reduction for graphics processing units enables genome-wide testing of epistasis in sporadic ALSabstractMOTIVATION: Epistasis, the presence of gene-gene interactions, has been hypothesized to be at the root of many common human diseases, but current genome-wide association studies largely ignore its role. Multifactor dimensionality reduction (MDR) is a powerful model-free method for detecting epistatic relationships between genes, but computational costs have made its application to genome-wide data difficult. Graphics processing units (GPUs), the hardware responsible for rendering computer games, are powerful parallel processors. Using GPUs to run MDR on a genome-wide dataset allows for statistically rigorous testing of epistasis. RESULTS: The implementation of MDR for GPUs (MDRGPU) includes core features of the widely used Java software package, MDR. This GPU implementation allows for large-scale analysis of epistasis at a dramatically lower cost than the standard CPU-based implementations. As a proof-of-concept, we applied this software to a genome-wide study of sporadic amyotrophic lateral sclerosis (ALS). We discovered a statistically significant two-SNP classifier and subsequently replicated the significance of these two SNPs in an independent study of ALS. MDRGPU makes the large-scale analysis of epistasis tractable and opens the door to statistically rigorous testing of interactions in genome-wide datasets. AVAILABILITY: MDRGPU is open source and available free of charge from http://www.sourceforge.net/projects/mdr. Casey S. Greene, Nicholas A. Sinnott-Armstrong, Daniel S. Himmelstein, Paul J. Park, Jason H. Moore, Brent T. Harris |
Bioinform. | 1 |
| 2009 | Nature-inspired algorithms for the genetic analysis of epistasis in common human diseases: Theoretical assessment of wrapper vs. filter approachesabstractIn human genetics, new technological methods allow researchers to collect a wealth of information about genetic variation among individuals quickly and relatively inexpensively. Studies examining more than one half of a million points of genetic variation are the new standard. Quickly analyzing these data to discover single gene effects is both feasible and often done. Unfortunately as our understanding of common human disease grows, we now believe it is likely that an individual's risk of these common diseases is not determined by simple single gene effects. Instead it seems likely that risk will be determined by nonlinear gene-gene interactions, also known as epistasis. Unfortunately searching for these nonlinear effects requires either effective search strategies or exhaustive search. Previously we have employed both filter and nature-inspired probabilistic search wrapper approaches such as genetic programming (GP) and ant colony optimization (ACO) to this problem. We have discovered that for this problem, expert knowledge is critical if we are to discover these interactions. Here we theoretically analyze both an expert knowledge filter and a simple expert-knowledge-aware wrapper. We show that under certain assumptions, the filter strategy leads to the highest power. Finally we discuss the implications of this work for this type of problem, and discuss how probabilistic search strategies which outperform a filtering approach may be designed. Casey S. Greene, Jeff Kiralis, Jason H. Moore |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | Sensible initialization using expert knowledge for genome-wide analysis of epistasis using genetic programmingabstractFor biomedical researchers it is now possible to measure large numbers of DNA sequence variations across the human genome. Measuring hundreds of thousands of variations is now routine, but single variations which consistently predict an individual's risk of common human disease have proven elusive. Instead of single variants determining the risk of common human diseases, it seems more likely that disease risk is best modeled by interactions between biological components. The evolutionary computing challenge now is to effectively explore interactions in these large datasets and identify combinations of variations which are robust predictors of common human diseases such as bladder cancer. One promising approach to this problem is genetic programming (GP). A GP approach for this problem will use darwinian inspired evolution to evolve programs which find and model attribute interactions which predict an individual's risk of common human diseases. The goal of this study is to develop and evaluate two initializers for this domain. We develop a probabilistic initializer which uses expert knowledge to select attributes and an enumerative initializer which maximizes attribute diversity in the generated population.We compare these initializers to a random initializer which displays no preference for attributes. We show that the expert-knowledge-aware probabilistic initializer significantly outperforms both the random initializer and the enumerative initializer.We discuss implications of these results for the design of GP strategies which are able to detect and characterize predictors of common human diseases. Casey S. Greene, Bill C. White, Jason H. Moore |
IEEE Congress on Evolutionary Computation | 1 |
| 2009 | Development and evaluation of an open-ended computational evolution system for the creation of digital organisms with complex genetic architectureabstractEpistasis, or gene-gene interaction, is a ubiquitous phenomenon that is inadequately addressed in human genetic studies. There are few tools that can accurately identify high-order epistatic interactions, and there is a lack of general understanding as to how epistatic interactions fit into genetic architecture. Here we approach both problems through the lens of genetic programming (GP). It has recently been proposed that increasing open-endedness of GP will result in more complex solutions that better acknowledge the complexity of human genetic datasets. Moreover, the solutions evolved in open-ended GP can serve as model organisms in which to study general effects of epistasis on phenotype. Here we introduce a prototype computational evolution system that implements an open-ended GP and generates organisms that display epistatic interactions. These interactions are significantly more prevalent and have a greater effect on fitness than epistatic interactions in organisms generated in the absence of selection. Anna L. Tyler, Bill C. White, Casey S. Greene, Paul C. Andrews, Richard Cowper-Sallari, Jason H. Moore |
IEEE Congress on Evolutionary Computation | 3 |
| 2009 | Environmental noise improves epistasis models of genetic data discovered using a computational evolution systemabstractCommon human diseases likely result from nonlinear interactions between multiple DNA sequence variations. One goal of human genetics is to use data mining and machine learning methods to identify combinations of genetic variations that are predictive of discrete measures of health in human population data. "Artificial evolution" approaches loosely based on real biological processes have been developed and applied in this domain, but it has recently been suggested that "computational evolution" approaches which incorporate additional biological and evolutionary complexity into existing algorithms will be more likely to solve problems of interest to biologists and biomedical researchers. Here we introduce a method to evolve compact solutions by adding environmental noise to a dataset during fitness evaluation. In ecological systems a highly specialized organism can fail to thrive as the environment changes. By introducing numerous small changes into training data, i.e. the environment, during evolution we similarly drive selection towards more general solutions. We show that this improves the power of the computational evolution system when modest amounts of noise are used. Furthermore, this method of changing the environment in which fitness is evaluated with small perturbations fits within the computational evolution framework and is an effective method of controlling solution size for problems where the data are likely to be noisy. Casey S. Greene, Douglas P. Hill, Jason H. Moore |
GECCO | 1 |
| 2008 | Using expert knowledge in initialization for genome-wide analysis of epistasis using genetic programmingabstractIn human genetics it is now possible to measure large numbers of DNA sequence variations across the human genome. Given current knowledge about biological networks and disease processes it seems likely that disease risk can best be modeled by interactions between biological components, which may be examined as interacting DNA sequence variations. The machine learning challenge is to e.ectively explore interactions in these datasets to identify combinations of variations which are predictive of common human diseases. Genetic programming is a promising approach to this problem. The goal of this study is to examine the role that an expert knowledge aware initializer can play in the framework of genetic programming. We show that this expert knowledge aware initializer outperforms both a random initializer and an enumerative initializer. Casey S. Greene, Bill C. White, Jason H. Moore |
GECCO | 1 |