Jason H. Moore

dblp:07/315 · DBLP profile ↗
← Back
129ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0002-5015-1099ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 69 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 62 · 7 first-author · 6 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Benchmarking large language models for identifying transcription factor regulatory interactions
abstract
MOTIVATION: Transcription factors (TFs) and their target genes form regulatory networks that control gene expression and influence diverse biological processes and disease outcomes. Although multiple computational methods and curated databases have been developed to identify TF-target interactions, they often require specialized expertise. Large language models (LLMs) chatbots offer a more accessible alternative for querying TF-target interactions. In this study, we benchmarked four prominent LLMs, Anthropic's Claude 3.5 Sonnet, Google's Gemini 1.0 Pro, OpenAI's GPT-4o, and Meta's Llama3 8b, using 8432 literature-curated human TF-target interactions. We examined four regulatory categories: bidirectional, ambiguous, self-regulated, and unidirectional interactions. RESULTS: Under single-turn queries, Claude 3.5 Sonnet and GPT-4o outperformed the others, with balanced accuracies reaching 50.0 ± 7.6% (GPT-4o, self-regulated) and 48.2 ± 1.0% (Claude 3.5 Sonnet, unidirectional). Zero-temperature settings generally enhanced reproducibility, and multi-turn prompting improved performance for most models, increasing Claude 3.5 Sonnet's accuracy on self-regulated pairs by 32.6%. Excluding TF-target pairs with all unknown regulation types also generally improved accuracy, with unidirectional regulation reaching near 70% balanced accuracy in some cases. We also benchmarked Anthropic's Claude 3.5 Sonnet, Google's Gemini 2.0 Flash, OpenAI's GPT-4o, and Meta's Llama3 using 5148 experimentally derived TF-target interactions. Claude 3.5 Sonnet consistently outperformed the other models across conditions. Our findings highlight that prompt engineering and strategic use of model parameters consistently influence LLM chatbots' performance on TF-target identifications. This study establishes a benchmarking framework and demonstrates the potential of pre-trained general-purpose LLMs to support regulatory biology research, especially for researchers without extensive computational expertise. AVAILABILITY AND IMPLEMENTATION: The literature-based TF-target interactions ground truth were obtained from TRRUST v2 human dataset (www.grnpedia.org/trrust). The experimental derived TF-target interactions ground truth were obtained from TFLink Home Sapiens small-scale interaction table (https://tflink.net/). Processed TF-target interactions data and the analytical pipeline has been compiled as an interactive Python notebook file and is available at https://github.com/pengpclab/LLM-TF-interactions.
Lake Noel, Yi-Wen Hsiao, Yimeng He, Andrew Hung, Xiaojiang Cui, Edward Ray, Jason H. Moore, Pei-Chen Peng, Xiuzhen Huang
Bioinform.7
2026 Survival-LCS: Rule-Based Survival Analysis without Proportional Hazard Assumptions
abstract
Survival analysis is widely utilized to model time-to-event data across biomedical, epidemiological, and engineering domains. However, traditional methods such as Cox regression impose strong assumptions, including proportional hazards, and are often inadequate for capturing complex, non-linear relationships in high-dimensional or heterogeneous data. Rule-based machine learning algorithms, such as "ExSTraCS," can interpretably model complex biomedical associations in classification tasks. This study extends ExSTraCS to the challenges of right-censored survival data to yield "Survival-LCS," the first rule-based survival analysis algorithm that additionally handles heterogeneous feature types and missing values and makes no assumptions about baseline hazard or survival time distributions. Survival-LCS is evaluated and compared across a variety of simulated genetic survival datasets covering distinct genetic architectures (additive, epistatic, heterogeneous, and univariate), censoring values, number of features, and survival distributions (random monotonic spline, Gamma, Gaussian, and Weibull) using Integrated Brier Scores and statistical significance testing. Results demonstrate that Survival-LCS reliably captures complex associations with survival outcomes, performing competitively with standard approaches, particularly in settings that challenge traditional models. This work highlights the capability of rule-based algorithms to be adapted as interpretable, assumption-free (i.e., data-driven) survival analysis methods, particularly in domains with higher-dimensional data or complex underlying associations.
Alexa A. Woodward, Harsh Bandhey, Jason H. Moore, Ryan J. Urbanowicz
ACM Trans. Evol. Learn. Optim.3
2025 Ethics vs. Regulation: Converging Frameworks for Trustworthy Human-Centered AI in Biomedical Research
abstract
The accelerating impact of AI in biomedical research is driving significant advances in precision medicine. As these systems increasingly shape health outcomes, the imperative to develop trustworthy, reliable, and ethically grounded AI becomes more pressing, particularly in addressing concerns related to data integrity, patient safety, and equitable outcomes. While the potential of AI to transform biomedical research is clear, its responsible integration depends on more than technological capability. Ensuring that these systems are aligned with societal values requires a dual commitment: the operationalization of ethical principles throughout the AI life cycle and the establishment of robust regulatory mechanisms. Ethics provides the normative vision for fairness, accountability, and human dignity, whereas regulation translates these ideals into enforceable standards. This paper explores the convergence of these domains as a necessary foundation for developing trustworthy human-centered AI in biomedical contexts. We provide practical guidance for AI developers and researchers on integrating proactive governance and translating ethical principles into actionable strategies to support equitable and responsible innovation.
Tayo Obafemi-Ajayi, Tiffani J. Bright, Emily F. Wong, Donald C. Wunsch II, Joan Peckham, Jason H. Moore
IJCNN6
2025 ESCARGOT: an AI agent leveraging large language models, dynamic graph of thoughts, and biomedical knowledge graphs for enhanced reasoning
abstract
MOTIVATION: LLMs like GPT-4, despite their advancements, often produce hallucinations and struggle with integrating external knowledge effectively. While Retrieval-Augmented Generation (RAG) attempts to address this by incorporating external information, it faces significant challenges such as context length limitations and imprecise vector similarity search. ESCARGOT aims to overcome these issues by combining LLMs with a dynamic Graph of Thoughts and biomedical knowledge graphs, improving output reliability, and reducing hallucinations. RESULT: ESCARGOT significantly outperforms industry-standard RAG methods, particularly in open-ended questions that demand high precision. ESCARGOT also offers greater transparency in its reasoning process, allowing for the vetting of both code and knowledge requests, in contrast to the black-box nature of LLM-only or RAG-based approaches. AVAILABILITY AND IMPLEMENTATION: ESCARGOT is available as a pip package and on GitHub at: https://github.com/EpistasisLab/ESCARGOT.
Nicholas Matsumoto, Hyunjun Choi, Jay Moran, Miguel E. Hernandez, Mythreye Venkatesan, Jui-Hsuan Chang, Zhiping Paul Wang, Jason H. Moore
Bioinform.9
2024 Survival-LCS: A Rule-Based Machine Learning Approach to Survival Analysis
Alexa A. Woodward, Harsh Bandhey, Jason H. Moore, Ryan J. Urbanowicz
GECCO3
2024 A review of feature selection strategies utilizing graph data structures and Knowledge Graphs
abstract
Feature selection in Knowledge Graphs (KGs) is increasingly utilized in diverse domains, including biomedical research, Natural Language Processing (NLP), and personalized recommendation systems. This paper delves into the methodologies for feature selection (FS) within KGs, emphasizing their roles in enhancing machine learning (ML) model efficacy, hypothesis generation, and interpretability. Through this comprehensive review, we aim to catalyze further innovation in FS for KGs, paving the way for more insightful, efficient, and interpretable analytical models across various domains. Our exploration reveals the critical importance of scalability, accuracy, and interpretability in FS techniques, advocating for the integration of domain knowledge to refine the selection process. We highlight the burgeoning potential of multi-objective optimization and interdisciplinary collaboration in advancing KG FS, underscoring the transformative impact of such methodologies on precision medicine, among other fields. The paper concludes by charting future directions, including the development of scalable, dynamic FS algorithms and the integration of explainable AI principles to foster transparency and trust in KG-driven models.
Pedro Henrique Ribeiro, Christina M. Ramirez, Jason H. Moore
Briefings Bioinform.4
2024 KRAGEN: a knowledge graph-enhanced RAG framework for biomedical problem solving using large language models
abstract
MOTIVATION: Answering and solving complex problems using a large language model (LLM) given a certain domain such as biomedicine is a challenging task that requires both factual consistency and logic, and LLMs often suffer from some major limitations, such as hallucinating false or irrelevant information, or being influenced by noisy data. These issues can compromise the trustworthiness, accuracy, and compliance of LLM-generated text and insights. RESULTS: Knowledge Retrieval Augmented Generation ENgine (KRAGEN) is a new tool that combines knowledge graphs, Retrieval Augmented Generation (RAG), and advanced prompting techniques to solve complex problems with natural language. KRAGEN converts knowledge graphs into a vector database and uses RAG to retrieve relevant facts from it. KRAGEN uses advanced prompting techniques: namely graph-of-thoughts (GoT), to dynamically break down a complex problem into smaller subproblems, and proceeds to solve each subproblem by using the relevant knowledge through the RAG framework, which limits the hallucinations, and finally, consolidates the subproblems and provides a solution. KRAGEN's graph visualization allows the user to interact with and evaluate the quality of the solution's GoT structure and logic. AVAILABILITY AND IMPLEMENTATION: KRAGEN is deployed by running its custom Docker containers. KRAGEN is available as open-source from GitHub at: https://github.com/EpistasisLab/KRAGEN.
Nicholas Matsumoto, Jay Moran, Hyunjun Choi, Miguel E. Hernandez, Mythreye Venkatesan, Zhiping Paul Wang, Jason H. Moore
Bioinform.7
2024 Interpretable deep clustering survival machines for Alzheimer's disease subtype discovery
Bojian Hou, Zixuan Wen, Jingxuan Bao, Richard Zhang 0001, Boning Tong, Shu Yang 0009, Junhao Wen 0002, Yuhan Cui, Jason H. Moore, Andrew J. Saykin, Heng Huang 0001, Paul M. Thompson, Marylyn D. Ritchie, Christos Davatzikos, Li Shen 0001
Medical Image Anal.9
2023 Faster Convergence with Lexicase Selection in Tree-Based Automated Machine Learning
Nicholas Matsumoto, Anil Kumar Saini, Pedro Henrique Ribeiro, Hyunjun Choi, Alena Orlenko, Leo-Pekka Lyytikäinen, Jari O. Laurikka, Terho Lehtimäki, Sandra Batista, Jason H. Moore
EuroGP10
2023 Aliro: an automated machine learning tool leveraging large language models
abstract
MOTIVATION: Biomedical and healthcare domains generate vast amounts of complex data that can be challenging to analyze using machine learning tools, especially for researchers without computer science training. RESULTS: Aliro is an open-source software package designed to automate machine learning analysis through a clean web interface. By infusing the power of large language models, the user can interact with their data by seamlessly retrieving and executing code pulled from the large language model, accelerating automated discovery of new insights from data. Aliro includes a pre-trained machine learning recommendation system that can assist the user to automate the selection of machine learning algorithms and its hyperparameters and provides visualization of the evaluated model and data. AVAILABILITY AND IMPLEMENTATION: Aliro is deployed by running its custom Docker containers. Aliro is available as open-source from GitHub at: https://github.com/EpistasisLab/Aliro.
Hyunjun Choi, Jay Moran, Nicholas Matsumoto, Miguel E. Hernandez, Jason H. Moore
Bioinform.5
2023 Ten simple rules for managing laboratory information
abstract
Information is the cornerstone of research, from experimental (meta)data and computational processes to complex inventories of reagents and equipment. These 10 simple rules discuss best practices for leveraging laboratory information management systems to transform this large information load into useful scientific findings.
Casey-Tyler Berezin, Luis U. Aguilera, Sonja Billerbeck, Philip E. Bourne, Douglas Densmore, Paul S. Freemont, Thomas E. Gorochowski, Sarah I. Hernandez, Nathan J. Hillson, Connor R. King, Michael Köpke, Shuyi Ma, Katie M. Miller, Tae Seok Moon, Jason H. Moore, Brian Munsky, Chris J. Myers, Dequina A. Nicholas, Samuel J. Peccoud, Jean Peccoud
PLoS Comput. Biol.15
2022 Preference Matrix Guided Sparse Canonical Correlation Analysis for Genetic Study of Quantitative Traits in Alzheimer's Disease
abstract
Investigating the relationship between genetic variation and phenotypic traits is a key issue in quantitative genetics. Specifically for Alzheimer's disease, the association between genetic markers and quantitative traits remains vague while, once identified, will provide valuable guidance for the study and development of genetic-based treatment approaches. Currently, to analyze the association of two modalities, sparse canonical correlation analysis (SCCA) is commonly used to compute one sparse linear combination of the variable features for each modality, giving a pair of linear combination vectors in total that maximizes the cross-correlation between the analyzed modalities. One drawback of the plain SCCA model is that the existing findings and knowledge cannot be integrated into the model as priors to help extract interesting correlation as well as identify biologically meaningful genetic and phenotypic markers. To bridge this gap, we introduce preference matrix guided SCCA (PM-SCCA) that not only takes priors encoded as a preference matrix but also maintains computational simplicity. A simulation study and a real-data experiment are conducted to investigate the effectiveness of the model. Both experiments demonstrate that the proposed PM-SCCA model can capture not only genotype-phenotype correlation but also relevant features effectively.
Jiahang Sha, Jingxuan Bao, Kefei Liu 0001, Shu Yang 0009, Zixuan Wen, Yuhan Cui, Junhao Wen 0002, Christos Davatzikos, Jason H. Moore, Andrew J. Saykin, Qi Long, Li Shen 0001
BIBM9
2022 A manifesto on explainability for artificial intelligence in medicine
abstract
The rapid increase of interest in, and use of, artificial intelligence (AI) in computer applications has raised a parallel concern about its ability (or lack thereof) to provide understandable, or explainable, output to users. This concern is especially legitimate in biomedical contexts, where patient safety is of paramount importance. This position paper brings together seven researchers working in the field with different roles and perspectives, to explore in depth the concept of explainable AI, or XAI, offering a functional definition and conceptual framework or model that can be used when considering XAI. This is followed by a series of desiderata for attaining explainability in AI, each of which touches upon a key domain in biomedicine.
Carlo Combi, Beatrice Amico, Riccardo Bellazzi, Andreas Holzinger, Jason H. Moore, Marinka Zitnik, John H. Holmes
Artif. Intell. Medicine5
2022 PMLB v1.0: an open-source dataset collection for benchmarking machine learning methods
abstract
MOTIVATION: Novel machine learning and statistical modeling studies rely on standardized comparisons to existing methods using well-studied benchmark datasets. Few tools exist that provide rapid access to many of these datasets through a standardized, user-friendly interface that integrates well with popular data science workflows. RESULTS: This release of PMLB (Penn Machine Learning Benchmarks) provides the largest collection of diverse, public benchmark datasets for evaluating new machine learning and data science methods aggregated in one location. v1.0 introduces a number of critical improvements developed following discussions with the open-source community. AVAILABILITY AND IMPLEMENTATION: PMLB is available at https://github.com/EpistasisLab/pmlb. Python and R interfaces for PMLB can be installed through the Python Package Index and Comprehensive R Archive Network, respectively.
Joseph D. Romano, Trang T. Le, William G. La Cava, John T. Gregg, Daniel J. Goldberg, Praneel Chakraborty, Natasha L. Ray, Daniel S. Himmelstein, Weixuan Fu, Jason H. Moore
Bioinform.10
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics31
2022 Multi-task learning based structured sparse canonical correlation analysis for brain imaging genetics
Mansu Kim, Eun Jeong Min, Kefei Liu 0001, Andrew J. Saykin, Jason H. Moore, Qi Long, Li Shen 0001
Medical Image Anal.6
2022 Spatial Pyramid Pooling With 3D Convolution Improves Lung Cancer Detection
abstract
Lung cancer is the leading cause of cancer deaths. Low-dose computed tomography (CT)screening has been shown to significantly reduce lung cancer mortality but suffers from a high false positive rate that leads to unnecessary diagnostic procedures. The development of deep learning techniques has the potential to help improve lung cancer screening technology. Here we present the algorithm, DeepScreener, which can predict a patient's cancer status from a volumetric lung CT scan. DeepScreener is based on our model of Spatial Pyramid Pooling, which ranked 16th of 1972 teams (top 1 percent)in the Data Science Bowl 2017 competition (DSB2017), evaluated with the challenge datasets. Here we test the algorithm with an independent set of 1449 low-dose CT scans of the National Lung Screening Trial (NLST)cohort, and we find that DeepScreener has consistent performance of high accuracy. Furthermore, by combining Spatial Pyramid Pooling and 3D Convolution, it achieves an AUC of 0.892, surpassing the previous state-of-the-art algorithms using only 3D convolution. The advancement of deep learning algorithms can potentially help improve lung cancer detection with low-dose CT scans.
Jason L. Causey, Xianghao Chen, Wei Dong 0003, Karl Walker, Jake A. Qualls, Jonathan W. Stubblefield, Jason H. Moore, Yuanfang Guan, Xiuzhen Huang
IEEE ACM Trans. Comput. Biol. Bioinform.8
2022 Genetic Analysis of Coronary Artery Disease Using Tree-Based Automated Machine Learning Informed By Biology-Based Feature Selection
abstract
Machine Learning (ML) approaches are increasingly being used in biomedical applications. Important challenges of ML include choosing the right algorithm and tuning the parameters for optimal performance. Automated ML (AutoML) methods, such as Tree-based Pipeline Optimization Tool (TPOT), have been developed to take some of the guesswork out of ML thus making this technology available to users from more diverse backgrounds. The goals of this study were to assess applicability of TPOT to genomics and to identify combinations of single nucleotide polymorphisms (SNPs) associated with coronary artery disease (CAD), with a focus on genes with high likelihood of being good CAD drug targets. We leveraged public functional genomic resources to group SNPs into biologically meaningful sets to be selected by TPOT. We applied this strategy to data from the U.K. Biobank, detecting a strikingly recurrent signal stemming from a group of 28 SNPs. Importance analysis of these SNPs uncovered functional relevance of the top SNPs to genes whose association with CAD is supported in the literature and other resources. Furthermore, we employed game-theory based metrics to study SNP contributions to individual-level TPOT predictions and discover distinct clusters of well-predicted CAD cases. The latter indicates a promising approach towards precision medicine.
Elisabetta Manduchi, Trang T. Le, Weixuan Fu, Jason H. Moore
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 Towards effective GP multi-class classification based on dynamic targets
abstract
In the multi-class classification problem GP plays an important role when combined with other non-GP classifiers. However, when GP performs the actual classification (without relying on other classifiers) its classification accuracy is low. This is especially true when the number of classes is high. In this paper, we present DTC, a GP classifier that leverages the effectiveness of the dynamic target approach to evolve a set of discriminant functions (one for each class). Notably, DTC is the first GP classifier that defines the fitness of individuals by using the synergistic combination of linear scaling and the hinge-loss function (commonly used by SVM). Differently, most previous GP classifiers use the number of correct classifications to drive the evolution. We compare DTC with eight state-of-art multi-class classification techniques (e.g., RF, RS, MLP, and SVM) on eight popular datasets. The results show that DTC achieves competitive classification accuracy even with 15 classes, without relying on other classifiers.
Stefano Ruberto, Valerio Terragni, Jason H. Moore
GECCO3
2021 Evaluating recommender systems for AI-driven biomedical informatics
abstract
MOTIVATION: Many researchers with domain expertise are unable to easily apply machine learning (ML) to their bioinformatics data due to a lack of ML and/or coding expertise. Methods that have been proposed thus far to automate ML mostly require programming experience as well as expert knowledge to tune and apply the algorithms correctly. Here, we study a method of automating biomedical data science using a web-based AI platform to recommend model choices and conduct experiments. We have two goals in mind: first, to make it easy to construct sophisticated models of biomedical processes; and second, to provide a fully automated AI agent that can choose and conduct promising experiments for the user, based on the user's experiments as well as prior knowledge. To validate this framework, we conduct an experiment on 165 classification problems, comparing to state-of-the-art, automated approaches. Finally, we use this tool to develop predictive models of septic shock in critical care patients. RESULTS: We find that matrix factorization-based recommendation systems outperform metalearning methods for automating ML. This result mirrors the results of earlier recommender systems research in other domains. The proposed AI is competitive with state-of-the-art automated ML methods in terms of choosing optimal algorithm configurations for datasets. In our application to prediction of septic shock, the AI-driven analysis produces a competent ML model (AUROC 0.85±0.02) that performs on par with state-of-the-art deep learning results for this task, with much less computational effort. AVAILABILITY AND IMPLEMENTATION: PennAI is available free of charge and open-source. It is distributed under the GNU public license (GPL) version 3. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
William G. La Cava, Heather Williams, Weixuan Fu, Steven Vitale, Durga Srivatsan, Jason H. Moore
Bioinform.6
2021 treeheatr: an R package for interpretable decision tree visualizations
abstract
SUMMARY: treeheatr is an R package for creating interpretable decision tree visualizations with the data represented as a heatmap at the tree's leaf nodes. The integrated presentation of the tree structure along with an overview of the data efficiently illustrates how the tree nodes split up the feature space and how well the tree model performs. This visualization can also be examined in depth to uncover the correlation structure in the data and importance of each feature in predicting the outcome. Implemented in an easily installed package with a detailed vignette, treeheatr can be a useful teaching tool to enhance students' understanding of a simple decision tree model before diving into more complex tree-based machine learning methods. AVAILABILITY AND IMPLEMENTATION: The treeheatr package is freely available under the permissive MIT license at https://trang1618.github.io/treeheatr and https://cran.r-project.org/package=treeheatr. It comes with a detailed vignette that is automatically built with GitHub Actions continuous integration.
Trang T. Le, Jason H. Moore
Bioinform.2
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.26
2021 Use of electronic health records to support a public health response to the COVID-19 pandemic in the United States: a perspective from 15 academic medical centers
abstract
Our goal is to summarize the collective experience of 15 organizations in dealing with uncoordinated efforts that result in unnecessary delays in understanding, predicting, preparing for, containing, and mitigating the COVID-19 pandemic in the US. Response efforts involve the collection and analysis of data corresponding to healthcare organizations, public health departments, socioeconomic indicators, as well as additional signals collected directly from individuals and communities. We focused on electronic health record (EHR) data, since EHRs can be leveraged and scaled to improve clinical care, research, and to inform public health decision-making. We outline the current challenges in the data ecosystem and the technology infrastructure that are relevant to COVID-19, as witnessed in our 15 institutions. The infrastructure includes registries and clinical data networks to support population-level analyses. We propose a specific set of strategic next steps to increase interoperability, overall organization, and efficiencies.
Subha Madhavan, Lisa Bastarache, Jeffrey S. Brown, Atul J. Butte, David A. Dorr, Peter J. Embí, Charles P. Friedman, Kevin B. Johnson, Jason H. Moore, Isaac S. Kohane, Philip R. O. Payne, Jessica D. Tenenbaum, Mark G. Weiner, Adam B. Wilcox, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.9
2020 Harnessing Electronic Health Records to Study Emerging Environmental Disasters: A Proof of Concept with PFAS
Mary Regina Boland, Lena M. Davidson, Silvia P. Canelón, Jessica R. Meeker, Trevor M. Penning, John H. Holmes, Jason H. Moore
AMIA7
2020 Explainable Artificial Intelligence (XAI): Current Approaches and Paths to the Future
John H. Holmes, Riccardo Bellazzi, Carlo Combi, Jason H. Moore, Niels Peek
AMIA4
2020 Benchmarking Manifold Learning Methods on a Large Collection of Datasets
Patryk Orzechowski, Franciszek Magiera, Jason H. Moore
EuroGP3
2020 SGP-DT: Semantic Genetic Programming Based on Dynamic Targets
Stefano Ruberto, Valerio Terragni, Jason H. Moore
EuroGP3
2020 Genetic programming approaches to learning fair classifiers
abstract
Society has come to rely on algorithms like classifiers for important decision making, giving rise to the need for ethical guarantees such as fairness. Fairness is typically defined by asking that some statistic of a classifier be approximately equal over protected groups within a population. In this paper, current approaches to fairness are discussed and used to motivate algorithmic proposals that incorporate fairness into genetic programming for classification. We propose two ideas. The first is to incorporate a fairness objective into multi-objective optimization. The second is to adapt lexicase selection to define cases dynamically over intersections of protected groups. We describe why lexicase selection is well suited to pressure models to perform well across the potentially infinitely many subgroups over which fairness is desired. We use a recent genetic programming approach to construct models on four datasets for which fairness constraints are necessary, and empirically compare performance to prior methods utilizing game-theoretic solutions. Methods are assessed based on their ability to generate trade-offs of subgroup fairness and accuracy that are Pareto optimal. The result show that genetic programming methods in general, and random search in particular, are well suited to this task.
William G. La Cava, Jason H. Moore
GECCO2
2020 Image Feature Learning with Genetic Programming
Stefano Ruberto, Valerio Terragni, Jason H. Moore
PPSN (2)3
2020 How Computational Experiments Can Improve Our Understanding of the Genetic Architecture of Common Human Diseases
abstract
Susceptibility to common human diseases such as cancer is influenced by many genetic and environmental factors that work together in a complex manner. The state of the art is to perform a genome-wide association study (GWAS) that measures millions of single-nucleotide polymorphisms (SNPs) throughout the genome followed by a one-SNP-at-a-time statistical analysis to detect univariate associations. This approach has identified thousands of genetic risk factors for hundreds of diseases. However, the genetic risk factors detected have very small effect sizes and collectively explain very little of the overall heritability of the disease. Nonetheless, it is assumed that the genetic component of risk is due to many independent risk factors that contribute additively. The fact that many genetic risk factors with small effects can be detected is taken as evidence to support this notion. It is our working hypothesis that the genetic architecture of common diseases is partly driven by non-additive interactions. To test this hypothesis, we developed a heuristic simulation-based method for conducting experiments about the complexity of genetic architecture. We show that a genetic architecture driven by complex interactions is highly consistent with the magnitude and distribution of univariate effects seen in real data. We compare our results with measures of univariate and interaction effects from two large-scale GWASs of sporadic breast cancer and find evidence to support our hypothesis that is consistent with the results of our computational experiment.
Jason H. Moore, Randal S. Olson, Peter Schmitt, Yong Chen 0016, Elisabetta Manduchi
Artif. Life1
2020 Scaling tree-based automated machine learning to biomedical big data with a feature set selector
abstract
MOTIVATION: Automated machine learning (AutoML) systems are helpful data science assistants designed to scan data for novel features, select appropriate supervised learning models and optimize their parameters. For this purpose, Tree-based Pipeline Optimization Tool (TPOT) was developed using strongly typed genetic programing (GP) to recommend an optimized analysis pipeline for the data scientist's prediction problem. However, like other AutoML systems, TPOT may reach computational resource limits when working on big data such as whole-genome expression data. RESULTS: We introduce two new features implemented in TPOT that helps increase the system's scalability: Feature Set Selector (FSS) and Template. FSS provides the option to specify subsets of the features as separate datasets, assuming the signals come from one or more of these specific data subsets. FSS increases TPOT's efficiency in application on big data by slicing the entire dataset into smaller sets of features and allowing GP to select the best subset in the final pipeline. Template enforces type constraints with strongly typed GP and enables the incorporation of FSS at the beginning of each pipeline. Consequently, FSS and Template help reduce TPOT computation time and may provide more interpretable results. Our simulations show TPOT-FSS significantly outperforms a tuned XGBoost model and standard TPOT implementation. We apply TPOT-FSS to real RNA-Seq data from a study of major depressive disorder. Independent of the previous study that identified significant association with depression severity of two modules, TPOT-FSS corroborates that one of the modules is largely predictive of the clinical diagnosis of each individual. AVAILABILITY AND IMPLEMENTATION: Detailed simulation and analysis code needed to reproduce the results in this study is available at https://github.com/lelaboratoire/tpot-fss. Implementation of the new TPOT operators is available at https://github.com/EpistasisLab/tpot. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Trang T. Le, Weixuan Fu, Jason H. Moore
Bioinform.3
2020 Model selection for metabolomics: predicting diagnosis of coronary artery disease using automated machine learning
abstract
MOTIVATION: Selecting the optimal machine learning (ML) model for a given dataset is often challenging. Automated ML (AutoML) has emerged as a powerful tool for enabling the automatic selection of ML methods and parameter settings for the prediction of biomedical endpoints. Here, we apply the tree-based pipeline optimization tool (TPOT) to predict angiographic diagnoses of coronary artery disease (CAD). With TPOT, ML models are represented as expression trees and optimal pipelines discovered using a stochastic search method called genetic programing. We provide some guidelines for TPOT-based ML pipeline selection and optimization-based on various clinical phenotypes and high-throughput metabolic profiles in the Angiography and Genes Study (ANGES). RESULTS: We analyzed nuclear magnetic resonance-derived lipoprotein and metabolite profiles in the ANGES cohort with a goal to identify the role of non-obstructive CAD patients in CAD diagnostics. We performed a comparative analysis of TPOT-generated ML pipelines with selected ML classifiers, optimized with a grid search approach, applied to two phenotypic CAD profiles. As a result, TPOT-generated ML pipelines that outperformed grid search optimized models across multiple performance metrics including balanced accuracy and area under the precision-recall curve. With the selected models, we demonstrated that the phenotypic profile that distinguishes non-obstructive CAD patients from no CAD patients is associated with higher precision, suggesting a discrepancy in the underlying processes between these phenotypes. AVAILABILITY AND IMPLEMENTATION: TPOT is freely available via http://epistasislab.github.io/tpot/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Alena Orlenko, Daniel Kofink, Leo-Pekka Lyytikäinen, Kjell Nikus, Pashupati P. Mishra, Pekka Kuukasjärvi, Pekka J. Karhunen, Mika Kähönen, Jari O. Laurikka, Terho Lehtimäki, Folkert W. Asselbergs, Jason H. Moore
Bioinform.12
2020 Regional imaging genetic enrichment analysis
abstract
MOTIVATION: Brain imaging genetics aims to reveal genetic effects on brain phenotypes, where most studies examine phenotypes defined on anatomical or functional regions of interest (ROIs) given their biologically meaningful interpretation and modest dimensionality compared with voxelwise approaches. Typical ROI-level measures used in these studies are summary statistics from voxelwise measures in the region, without making full use of individual voxel signals. RESULTS: In this article, we propose a flexible and powerful framework for mining regional imaging genetic associations via voxelwise enrichment analysis, which embraces the collective effect of weak voxel-level signals and integrates brain anatomical annotation information. Our proposed method achieves three goals at the same time: (i) increase the statistical power by substantially reducing the burden of multiple comparison correction; (ii) employ brain annotation information to enable biologically meaningful interpretation and (iii) make full use of fine-grained voxelwise signals. We demonstrate our method on an imaging genetic analysis using data from the Alzheimer's Disease Neuroimaging Initiative, where we assess the collective regional genetic effects of voxelwise FDG-positron emission tomography measures between 116 ROIs and 565 373 single-nucleotide polymorphisms. Compared with traditional ROI-wise and voxelwise approaches, our method identified 2946 novel imaging genetic associations in addition to 33 ones overlapping with the two benchmark methods. In particular, two newly reported variants were further supported by transcriptome evidences from region-specific expression analysis. This demonstrates the promise of the proposed method as a flexible and powerful framework for exploring imaging genetic effects on the brain. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at https://github.com/lshen/RIGEA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaohui Yao, Shan Cong, Shannon L. Risacher, Andrew J. Saykin, Jason H. Moore, Li Shen 0001
Bioinform.6
2020 Embedding covariate adjustments in tree-based automated machine learning for biomedical big data analyses
abstract
BACKGROUND: A typical task in bioinformatics consists of identifying which features are associated with a target outcome of interest and building a predictive model. Automated machine learning (AutoML) systems such as the Tree-based Pipeline Optimization Tool (TPOT) constitute an appealing approach to this end. However, in biomedical data, there are often baseline characteristics of the subjects in a study or batch effects that need to be adjusted for in order to better isolate the effects of the features of interest on the target. Thus, the ability to perform covariate adjustments becomes particularly important for applications of AutoML to biomedical big data analysis. RESULTS: We developed an approach to adjust for covariates affecting features and/or target in TPOT. Our approach is based on regressing out the covariates in a manner that avoids 'leakage' during the cross-validation training procedure. We describe applications of this approach to toxicogenomics and schizophrenia gene expression data sets. The TPOT extensions discussed in this work are available at https://github.com/EpistasisLab/tpot/tree/v0.11.1-resAdj . CONCLUSIONS: In this work, we address an important need in the context of AutoML, which is particularly crucial for applications to bioinformatics and medical informatics, namely covariate adjustments. To this end we present a substantial extension of TPOT, a genetic programming based AutoML approach. We show the utility of this extension by applications to large toxicogenomics and differential gene expression data. The method is generally applicable in many other scenarios from the biomedical field.
Elisabetta Manduchi, Weixuan Fu, Joseph D. Romano, Stefano Ruberto, Jason H. Moore
BMC Bioinform.5
2020 Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithm
abstract
OBJECTIVES: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites. MATERIALS AND METHODS: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard). RESULTS: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was <3%, and the ratio of standard errors was <1.25 for all scenarios. ODAL2 achieved higher accuracy (with relative bias <0.1% and ratio of standard errors <1.05). In real data analysis, we investigated the associations of 100 medications with fetal loss during pregnancy. We found that ODAL1 provided estimates with relative bias <10% for 85% of medications, and ODAL2 has relative bias <10% for 99% of medications. For communication cost, ODAL1 requires transferring p numbers from each site to the local site and ODAL2 requires transferring (p×p+p) numbers from each site to the local site, where p is the number of parameters in the regression model. CONCLUSIONS: This study demonstrates that ODAL is privacy-preserving and communication-efficient with small bias and high statistical efficiency.
Rui Duan 0004, Mary Regina Boland, Howard H. Chang, Hua Xu 0001, Haitao Chu, Christopher H. Schmid, Christopher B. Forrest, John H. Holmes, Martijn J. Schuemie, Jesse A. Berlin, Jason H. Moore, Yong Chen 0016
J. Am. Medical Informatics Assoc.13
2020 Learning from local to global: An efficient distributed algorithm for modeling time-to-event data
abstract
OBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner.
Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016
J. Am. Medical Informatics Assoc.14
2020 An augmented estimation procedure for EHR-based association studies accounting for differential misclassification
abstract
OBJECTIVES: The ability to identify novel risk factors for health outcomes is a key strength of electronic health record (EHR)-based research. However, the validity of such studies is limited by error in EHR-derived phenotypes. The objective of this study was to develop a novel procedure for reducing bias in estimated associations between risk factors and phenotypes in EHR data. MATERIALS AND METHODS: The proposed method combines the strengths of a gold-standard phenotype obtained through manual chart review for a small validation set of patients and an automatically-derived phenotype that is available for all patients but is potentially error-prone (hereafter referred to as the algorithm-derived phenotype). An augmented estimator of associations is obtained by optimally combining these 2 phenotypes. We conducted simulation studies to evaluate the performance of the augmented estimator and conducted an analysis of risk factors for second breast cancer events using data on a cohort from Kaiser Permanente Washington. RESULTS: The proposed method was shown to reduce bias relative to an estimator using only the algorithm-derived phenotype and reduce variance compared to an estimator using only the validation data. DISCUSSION: Our simulation studies and real data application demonstrate that, compared to the estimator using validation data only, the augmented estimator has lower variance (ie, higher statistical efficiency). Compared to the estimator using error-prone EHR-derived phenotypes, the augmented estimator has smaller bias. CONCLUSIONS: The proposed estimator can effectively combine an error-prone phenotype with gold-standard data from a limited chart review in order to improve analyses of risk factors using EHR data.
Jiayi Tong, Jing Huang 0021, Jessica Chubak, Jason H. Moore, Rebecca A. Hubbard, Yong Chen 0016
J. Am. Medical Informatics Assoc.5
2020 A maximum likelihood approach to electronic health record phenotyping using positive and unlabeled patients
abstract
OBJECTIVE: Phenotyping patients using electronic health record (EHR) data conventionally requires labeled cases and controls. Assigning labels requires manual medical chart review and therefore is labor intensive. For some phenotypes, identifying gold-standard controls is prohibitive. We developed an accurate EHR phenotyping approach that does not require labeled controls. MATERIALS AND METHODS: Our framework relies on a random subset of cases, which can be specified using an anchor variable that has excellent positive predictive value and sensitivity independent of predictors. We proposed a maximum likelihood approach that efficiently leverages data from the specified cases and unlabeled patients to develop logistic regression phenotyping models, and compare model performance with existing algorithms. RESULTS: Our method outperformed the existing algorithms on predictive accuracy in Monte Carlo simulation studies, application to identify hypertension patients with hypokalemia requiring oral supplementation using a simulated anchor, and application to identify primary aldosteronism patients using real-world cases and anchor variables. Our method additionally generated consistent estimates of 2 important parameters, phenotype prevalence and the proportion of true cases that are labeled. DISCUSSION: Upon identification of an anchor variable that is scalable and transferable to different practices, our approach should facilitate development of scalable, transferable, and practice-specific phenotyping models. CONCLUSIONS: Our proposed approach enables accurate semiautomated EHR phenotyping with minimal manual labeling and therefore should greatly facilitate EHR clinical decision support and research.
Lingjiao Zhang, Xiruo Ding, Yanyuan Ma, Naveen Muthu, Imran Ajmal, Jason H. Moore, Daniel S. Herman
J. Am. Medical Informatics Assoc.6
2020 Ten simple rules for writing a paper about scientific software
abstract
Papers describing software are an important part of computational fields of scientific research. These "software papers" are unique in a number of ways, and they require special consideration to improve their impact on the scientific community and their efficacy at conveying important information. Here, we discuss 10 specific rules for writing software papers, covering some of the different scenarios and publication types that might be encountered, and important questions from which all computational researchers would benefit by asking along the way.
Joseph D. Romano, Jason H. Moore
PLoS Comput. Biol.2
2020 Gamorithm
abstract
Examining games from a fresh perspective, we present the idea of game-inspired and game-based algorithms, dubbed gamorithms.
Moshe Sipper, Jason H. Moore
IEEE Trans. Games2
2019 Interpretation of machine learning predictions for patient outcomes in electronic health records
William G. La Cava, Christopher R. Bauer, Jason H. Moore, Sarah A. Pendergrass
AMIA3
2019 Solution and Fitness Evolution (SAFE): A Study of Multiobjective Problems
abstract
We have recently presented SAFE-Solution And Fitness Evolution-a commensalistic coevolutionary algorithm that maintains two coevolving populations: a population of candidate solutions and a population of candidate objective functions. We showed that SAFE was successful at evolving solutions within a robotic maze domain. Herein we present an investigation of SAFE's adaptation and application to multiobjective problems, wherein candidate objective functions explore different weightings of each objective. Though preliminary, the results suggest that SAFE, and the concept of coevolving solutions and objective functions, can identify a similar set of optimal multiobjective solutions without explicitly employing a Pareto front for fitness calculation and parent selection. These findings support our hypothesis that the SAFE algorithm concept can not only solve complex problems, but can adapt to the challenge of problems with multiple objectives.
Moshe Sipper, Jason H. Moore, Ryan J. Urbanowicz
CEC2
2019 Solution and Fitness Evolution (SAFE): Coevolving Solutions and Their Objective Functions
Moshe Sipper, Jason H. Moore, Ryan J. Urbanowicz
EuroGP2
2019 Semantic variation operators for multidimensional genetic programming
abstract
Multidimensional genetic programming represents candidate solutions as sets of programs, and thereby provides an interesting framework for exploiting building block identification. Towards this goal, we investigate the use of machine learning as a way to bias which components of programs are promoted, and propose two semantic operators to choose where useful building blocks are placed during crossover. A forward stagewise crossover operator we propose leads to significant improvements on a set of regression problems, and produces state-of-the-art results in a large benchmark study. We discuss this architecture and others in terms of their propensity for allowing heuristic search to utilize information during the evolutionary process. Finally, we look at the collinearity and complexity of the data representations that result from these architectures, with a view towards disentangling factors of variation in application.
William G. La Cava, Jason H. Moore
GECCO2
2019 Learning concise representations for regression by evolving networks of trees
William G. La Cava, Tilak Raj Singh, James Taggart, Srinivas Suri, Jason H. Moore
ICLR (Poster)5
2019 STatistical Inference Relief (STIR) feature selection
abstract
MOTIVATION: Relief is a family of machine learning algorithms that uses nearest-neighbors to select features whose association with an outcome may be due to epistasis or statistical interactions with other features in high-dimensional data. Relief-based estimators are non-parametric in the statistical sense that they do not have a parameterized model with an underlying probability distribution for the estimator, making it difficult to determine the statistical significance of Relief-based attribute estimates. Thus, a statistical inferential formalism is needed to avoid imposing arbitrary thresholds to select the most important features. We reconceptualize the Relief-based feature selection algorithm to create a new family of STatistical Inference Relief (STIR) estimators that retains the ability to identify interactions while incorporating sample variance of the nearest neighbor distances into the attribute importance estimation. This variance permits the calculation of statistical significance of features and adjustment for multiple testing of Relief-based scores. Specifically, we develop a pseudo t-test version of Relief-based algorithms for case-control data. RESULTS: We demonstrate the statistical power and control of type I error of the STIR family of feature selection methods on a panel of simulated data that exhibits properties reflected in real gene expression data, including main effects and network interaction effects. We compare the performance of STIR when the adaptive radius method is used as the nearest neighbor constructor with STIR when the fixed-k nearest neighbor constructor is used. We apply STIR to real RNA-Seq data from a study of major depressive disorder and discuss STIR's straightforward extension to genome-wide association studies. AVAILABILITY AND IMPLEMENTATION: Code and data available at http://insilico.utulsa.edu/software/STIR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Trang T. Le, Ryan J. Urbanowicz, Jason H. Moore, Brett A. McKinney
Bioinform.3
2019 EBIC: an open source software for high-dimensional and big data analyses
abstract
MOTIVATION: In this paper, we present an open source package with the latest release of Evolutionary-based BIClustering (EBIC), a next-generation biclustering algorithm for mining genetic data. The major contribution of this paper is adding a full support for multiple graphics processing units (GPUs) support, which makes it possible to run efficiently large genomic data mining analyses. Multiple enhancements to the first release of the algorithm include integration with R and Bioconductor, and an option to exclude missing values from the analysis. RESULTS: Evolutionary-based BIClustering was applied to datasets of different sizes, including a large DNA methylation dataset with 436 444 rows. For the largest dataset we observed over 6.6-fold speedup in computation time on a cluster of eight GPUs compared to running the method on a single GPU. This proves high scalability of the method. AVAILABILITY AND IMPLEMENTATION: The latest version of EBIC could be downloaded from http://github.com/EpistasisLab/ebic. Installation and usage instructions are also available online. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Patryk Orzechowski, Jason H. Moore
Bioinform.2
2019 A Probabilistic and Multi-Objective Analysis of Lexicase Selection and ε-Lexicase Selection
abstract
Lexicase selection is a parent selection method that considers training cases individually, rather than in aggregate, when performing parent selection. Whereas previous work has demonstrated the ability of lexicase selection to solve difficult problems in program synthesis and symbolic regression, the central goal of this article is to develop the theoretical underpinnings that explain its performance. To this end, we derive an analytical formula that gives the expected probabilities of selection under lexicase selection, given a population and its behavior. In addition, we expand upon the relation of lexicase selection to many-objective optimization methods to describe the behavior of lexicase selection, which is to select individuals on the boundaries of Pareto fronts in high-dimensional space. We show analytically why lexicase selection performs more poorly for certain sizes of population and training cases, and show why it has been shown to perform more poorly in continuous error spaces. To address this last concern, we propose new variants of [Formula: see text]-lexicase selection, a method that modifies the pass condition in lexicase selection to allow near-elite individuals to pass cases, thereby improving selection performance with continuous errors. We show that [Formula: see text]-lexicase outperforms several diversity–maintenance strategies on a number of real-world and synthetic regression problems.
William G. La Cava, Thomas Helmuth, Lee Spector, Jason H. Moore
Evol. Comput.4
2019 Integration of genetic and clinical information to improve imputation of data missing from electronic health records
abstract
OBJECTIVE: Clinical data of patients' measurements and treatment history stored in electronic health record (EHR) systems are starting to be mined for better treatment options and disease associations. A primary challenge associated with utilizing EHR data is the considerable amount of missing data. Failure to address this issue can introduce significant bias in EHR-based research. Currently, imputation methods rely on correlations among the structured phenotype variables in the EHR. However, genetic studies have shown that many EHR-based phenotypes have a heritable component, suggesting that measured genetic variants might be useful for imputing missing data. In this article, we developed a computational model that incorporates patients' genetic information to perform EHR data imputation. MATERIALS AND METHODS: We used the individual single nucleotide polymorphism's association with phenotype variables in the EHR as input to construct a genetic risk score that quantifies the genetic contribution to the phenotype. Multiple approaches to constructing the genetic risk score were evaluated for optimal performance. The genetic score, along with phenotype correlation, is then used as a predictor to impute the missing values. RESULTS: To demonstrate the method performance, we applied our model to impute missing cardiovascular related measurements including low-density lipoprotein, heart failure, and aortic aneurysm disease in the electronic Medical Records and Genomics data. The integration method improved imputation's area-under-the-curve for binary phenotypes and decreased root-mean-square error for continuous phenotypes. CONCLUSION: Compared with standard imputation approaches, incorporating genetic information offers a novel approach that can utilize more of the EHR data for better performance in missing data imputation.
Ruowang Li, Yong Chen 0016, Jason H. Moore
J. Am. Medical Informatics Assoc.3
2019 A regression framework to uncover pleiotropy in large-scale electronic health record data
abstract
OBJECTIVE: Pleiotropy, where 1 genetic locus affects multiple phenotypes, can offer significant insights in understanding the complex genotype-phenotype relationship. Although individual genotype-phenotype associations have been thoroughly explored, seemingly unrelated phenotypes can be connected genetically through common pleiotropic loci or genes. However, current analyses of pleiotropy have been challenged by both methodologic limitations and a lack of available suitable data sources. MATERIALS AND METHODS: In this study, we propose to utilize a new regression framework, reduced rank regression, to simultaneously analyze multiple phenotypes and genotypes to detect pleiotropic effects. We used a large-scale biobank linked electronic health record data from the Penn Medicine BioBank to select 5 cardiovascular diseases (hypertension, cardiac dysrhythmias, ischemic heart disease, congestive heart failure, and heart valve disorders) and 5 mental disorders (mood disorders; anxiety, phobic and dissociative disorders; alcohol-related disorders; neurological disorders; and delirium dementia) to validate our framework. RESULTS: Compared with existing methods, reduced rank regression showed a higher power to distinguish known associated single-nucleotide polymorphisms from random single-nucleotide polymorphisms. In addition, genome-wide gene-based investigation of pleiotropy showed that reduced rank regression was able to identify candidate genetic variants with novel pleiotropic effects compared to existing methods. CONCLUSION: The proposed regression framework offers a new approach to account for the phenotype and genotype correlations when identifying pleiotropic effects. By jointly modeling multiple phenotypes and genotypes together, the method has the potential to distinguish confounding from causal genotype and phenotype associations.
Ruowang Li, Rui Duan 0004, Daniel J. Rader, Scott M. Damrauer, Jason H. Moore, Yong Chen 0016
J. Am. Medical Informatics Assoc.5
2018 Comparing adverse effects of Hepatitis C drugs using FAERS data
Jing Huang 0021, Xinyuan Zhang 0003, Jiayi Tong, Jingcheng Du, Rui Duan 0004, Liu Yang 0026, Jason H. Moore, Yong Chen 0016, Cui Tao
BIBM7
2018 Where are we now?: a large benchmark study of recent symbolic regression methods
abstract
In this paper we provide a broad benchmarking of recent genetic programming approaches to symbolic regression in the context of state of the art machine learning approaches. We use a set of nearly 100 regression benchmark problems culled from open source repositories across the web. We conduct a rigorous benchmarking of four recent symbolic regression approaches as well as nine machine learning approaches from scikit-learn. The results suggest that symbolic regression performs strongly compared to state-of-the-art gradient boosting algorithms, although in terms of running times is among the slowest of the available methodologies. We discuss the results in detail and point to future research directions that may allow symbolic regression to gain wider adoption in the machine learning community.
Patryk Orzechowski, William G. La Cava, Jason H. Moore
GECCO3
2018 Attribute tracking: strategies towards improved detection and characterization of complex associations
abstract
The detection, modeling and characterization of complex patterns of association in bioinformatics has focused on feature interactions and, more recently, instance-subgroup specific associations (e.g. genetic heterogeneity). Previously, attribute tracking was proposed as an instance-linked memory approach, leveraging the incremental learning of learning classifier systems (LCSs) to track which features were most useful in making class predictions within individual instances. These 'attribute tracking' signatures could later be used to characterize patterns of association in the data. While effective, true underlying patterns remain difficult to characterize in noisy problems, and the original approach places equal weight on tracked feature scores obtained early as well as late in learning. In this work we investigate alternative strategies for attribute tracking scoring, including the adoption of a time recency update scheme taken from reinforcement learning, to gain insight into how to optimize this approach to improve modeling performance and downstream pattern interpretability. We report mixed results over a variety of performance metrics that point to promising future directions for building effective building blocks and improving model interpretability.
Ryan J. Urbanowicz, Christopher Lo, John H. Holmes, Jason H. Moore
GECCO4
2018 runibic: a Bioconductor package for parallel row-based biclustering of gene expression data
abstract
Motivation: Biclustering is an unsupervised technique of simultaneous clustering of rows and columns of input matrix. With multiple biclustering algorithms proposed, UniBic remains one of the most accurate methods developed so far. Results: In this paper we introduce a Bioconductor package called runibic with parallel implementation of UniBic. For the convenience the algorithm was reimplemented, parallelized and wrapped within an R package called runibic. The package includes: (i) a couple of times faster parallel version of the original sequential algorithm, (ii) much more efficient memory management, (iii) modularity which allows to build new methods on top of the provided one and (iv) integration with the modern Bioconductor packages such as SummarizedExperiment, ExpressionSet and biclust. Availability and implementation: The package is implemented in R and is available from Bioconductor (starting from version 3.6) at the following URL http://bioconductor.org/packages/runibic with installation instructions and tutorial. Supplementary information: Supplementary data are available at Bioinformatics online.
Patryk Orzechowski, Artur Panszczyk, Xiuzhen Huang, Jason H. Moore
Bioinform.4
2018 EBIC: an evolutionary-based parallel biclustering algorithm for pattern discovery
abstract
Motivation: Biclustering algorithms are commonly used for gene expression data analysis. However, accurate identification of meaningful structures is very challenging and state-of-the-art methods are incapable of discovering with high accuracy different patterns of high biological relevance. Results: In this paper, a novel biclustering algorithm based on evolutionary computation, a sub-field of artificial intelligence, is introduced. The method called EBIC aims to detect order-preserving patterns in complex data. EBIC is capable of discovering multiple complex patterns with unprecedented accuracy in real gene expression datasets. It is also one of the very few biclustering methods designed for parallel environments with multiple graphics processing units. We demonstrate that EBIC greatly outperforms state-of-the-art biclustering methods, in terms of recovery and relevance, on both synthetic and genetic datasets. EBIC also yields results over 12 times faster than the most accurate reference algorithms. Availability and implementation: EBIC source code is available on GitHub at https://github.com/EpistasisLab/ebic. Supplementary information: Supplementary data are available at Bioinformatics online.
Patryk Orzechowski, Moshe Sipper, Xiuzhen Huang, Jason H. Moore
Bioinform.4
2018 PIE: A prior knowledge guided integrated likelihood estimation method for bias reduction in association studies using electronic health records data
abstract
OBJECTIVES: This study proposes a novel Prior knowledge guided Integrated likelihood Estimation (PIE) method to correct bias in estimations of associations due to misclassification of electronic health record (EHR)-derived binary phenotypes, and evaluates the performance of the proposed method by comparing it to 2 methods in common practice. METHODS: We conducted simulation studies and data analysis of real EHR-derived data on diabetes from Kaiser Permanente Washington to compare the estimation bias of associations using the proposed method, the method ignoring phenotyping errors, the maximum likelihood method with misspecified sensitivity and specificity, and the maximum likelihood method with correctly specified sensitivity and specificity (gold standard). The proposed method effectively leverages available information on phenotyping accuracy to construct a prior distribution for sensitivity and specificity, and incorporates this prior information through the integrated likelihood for bias reduction. RESULTS: Our simulation studies and real data application demonstrated that the proposed method effectively reduces the estimation bias compared to the 2 current methods. It performed almost as well as the gold standard method when the prior had highest density around true sensitivity and specificity. The analysis of EHR data from Kaiser Permanente Washington showed that the estimated associations from PIE were very close to the estimates from the gold standard method and reduced bias by 60%-100% compared to the 2 commonly used methods in current practice for EHR data. CONCLUSIONS: This study demonstrates that the proposed method can effectively reduce estimation bias caused by imperfect phenotyping in EHR-derived data by incorporating prior information through integrated likelihood.
Jing Huang 0021, Rui Duan 0004, Rebecca A. Hubbard, Yonghui Wu 0001, Jason H. Moore, Hua Xu 0001, Yong Chen 0016
J. Am. Medical Informatics Assoc.5
2018 Medication class enrichment analysis: a novel algorithm to analyze multiple pharmacologic exposures simultaneously using electronic health record data
abstract
Objective: Observational studies analyzing multiple exposures simultaneously have been limited by difficulty distinguishing relevant results from chance associations due to poor specificity. Set-based methods have been successfully used in genomics to improve signal-to-noise ratio. We present and demonstrate medication class enrichment analysis (MCEA), a signal-to-noise enhancement algorithm for observational data inspired by set-based methods. Materials and Methods: We used The Health Improvement Network database to study medications associated with Clostridium difficile infection (CDI). We performed case-control studies for each medication in The Health Improvement Network to obtain odds ratios (ORs) for association with CDI. We then calculated the association of each pharmacologic class with CDI using logistic regression and MCEA. We also performed simulation studies in which we assessed the sensitivity and specificity of logistic regression compared to MCEA for ORs 0.1-2.0. Results: When analyzing pharmacologic classes using logistic regression, 47 of 110 pharmacologic classes were identified as associated with CDI. When analyzing pharmacologic classes using MCEA, only fluoroquinolones, a class of antibiotics with biologically confirmed causation, and heparin products were associated with CDI. In simulation, MCEA had superior specificity compared to logistic regression across all tested effect sizes and equal or better sensitivity for all effect sizes besides those close to null. Discussion: Although these results demonstrate the promise of MCEA, additional studies that include inpatient administered medications are necessary for validation of the algorithm. Conclusions: In clinical and simulation studies, MCEA demonstrated superior sensitivity and specificity for identifying pharmacologic classes associated with CDI compared to logistic regression.
Ravy K. Vajravelu, Frank I. Scott, Ronac Mamtani, Hongzhe Li, Jason H. Moore, James D. Lewis
J. Am. Medical Informatics Assoc.5
2018 Relief-based feature selection: Introduction and review
Ryan J. Urbanowicz, Melissa Meeker, William G. La Cava, Randal S. Olson, Jason H. Moore
J. Biomed. Informatics5
2018 Benchmarking relief-based feature selection methods for bioinformatics data mining
Ryan J. Urbanowicz, Randal S. Olson, Peter Schmitt, Melissa Meeker, Jason H. Moore
J. Biomed. Informatics5
2018 Eleven quick tips for architecting biomedical informatics workflows with cloud computing
abstract
Cloud computing has revolutionized the development and operations of hardware and software across diverse technological arenas, yet academic biomedical research has lagged behind despite the numerous and weighty advantages that cloud computing offers. Biomedical researchers who embrace cloud computing can reap rewards in cost reduction, decreased development and maintenance workload, increased reproducibility, ease of sharing data and software, enhanced security, horizontal and vertical scalability, high availability, a thriving technology partner ecosystem, and much more. Despite these advantages that cloud-based workflows offer, the majority of scientific software developed in academia does not utilize cloud computing and must be migrated to the cloud by the user. In this article, we present 11 quick tips for architecting biomedical informatics workflows on compute clouds, distilling knowledge gained from experience developing, operating, maintaining, and distributing software and virtualized appliances on the world's largest cloud. Researchers who follow these tips stand to benefit immediately by migrating their workflows to cloud computing and embracing the paradigm of abstraction.
Brian S. Cole, Jason H. Moore
PLoS Comput. Biol.2
2017 A General Feature Engineering Wrapper for Machine Learning Using \epsilon -Lexicase Survival
William G. La Cava, Jason H. Moore
EuroGP2
2017 Genetic Programming Representations for Multi-dimensional Feature Learning in Biomedical Classification
William G. La Cava, Sara Silva, Leonardo Vanneschi, Lee Spector, Jason H. Moore
EvoApplications (1)5
2017 EVE: Cloud-Based Annotation of Human Genetic Variants
Brian S. Cole, Jason H. Moore
EvoApplications (1)2
2017 Improving the Reproducibility of Genetic Association Results Using Genotype Resampling Methods
Elizabeth R. Piette, Jason H. Moore
EvoApplications (1)2
2017 Ensemble representation learning: an analysis of fitness and survival for wrapper-based genetic programming methods
abstract
Recently we proposed a general, ensemble-based feature engineering wrapper (FEW) that was paired with a number of machine learning methods to solve regression problems. Here, we adapt FEW for supervised classification and perform a thorough analysis of fitness and survival methods within this framework. Our tests demonstrate that two fitness metrics, one introduced as an adaptation of the silhouette score, outperform the more commonly used Fisher criterion. We analyze survival methods and demonstrate that ϵ-lexicase survival works best across our test problems, followed by random survival which outperforms both tournament and deterministic crowding. We conduct a benchmark comparison to several classification methods using a large set of problems and show that FEW can improve the best classifier performance in several cases. We show that FEW generates consistent, meaningful features for a biomedical problem with different ML pairings.
William G. La Cava, Jason H. Moore
GECCO2
2017 Toward the automated analysis of complex diseases in genome-wide association studies using genetic programming
abstract
Machine learning has been gaining traction in recent years to meet the demand for tools that can efficiently analyze and make sense of the ever-growing databases of biomedical data in health care systems around the world. However, effectively using machine learning methods requires considerable domain expertise, which can be a barrier of entry for bioinformaticians new to computational data science methods. Therefore, off-the-shelf tools that make machine learning more accessible can prove invaluable for bioinformaticians. To this end, we have developed an open source pipeline optimization tool (TPOT-MDR) that uses genetic programming to automatically design machine learning pipelines for bioinformatics studies. In TPOT-MDR, we implement Multifactor Dimensionality Reduction (MDR) as a feature construction method for modeling higher-order feature interactions, and combine it with a new expert knowledge-guided feature selector for large biomedical data sets. We demonstrate TPOT-MDR's capabilities using a combination of simulated and real world data sets from human genetics and find that TPOT-MDR significantly outperforms modern machine learning methods such as logistic regression and eXtreme Gradient Boosting (XGBoost). We further analyze the best pipeline discovered by TPOT-MDR for a real world problem and highlight TPOT-MDR's ability to produce a high-accuracy solution that is also easily interpretable.
Andrew Sohn, Randal S. Olson, Jason H. Moore
GECCO3
2017 Network-based genome wide study of hippocampal imaging phenotype in Alzheimer's Disease to identify functional interaction modules
abstract
Identification of functional modules from biological network is a promising approach to enhance the statistical power of genome-wide association study (GWAS) and improve biological interpretation for complex diseases. The precise functions of genes are highly relevant to tissue context, while a majority of module identification studies are based on tissue-free biological networks that lacks phenotypic specificity. In this study, we propose a module identification method that maps the GWAS results of an imaging phenotype onto the corresponding tissue-specific functional interaction network by applying a machine learning framework. Ridge regression and support vector machine (SVM) models are constructed to re-prioritize GWAS results, followed by exploring hippocampus-relevant modules based on top predictions using GWAS top findings. We also propose a GWAS top-neighbor-based module identification approach and compare it with Ridge and SVM based approaches. Modules conserving both tissue specificity and GWAS discoveries are identified, showing the promise of the proposal method for providing insight into the mechanism of complex diseases.
Xiaohui Yao, Shannon L. Risacher, Jason H. Moore, Andrew J. Saykin, Li Shen 0001
ICASSP4
2017 Tissue-specific network-based genome wide study of amygdala imaging phenotypes to identify functional interaction modules
abstract
MOTIVATION: Network-based genome-wide association studies (GWAS) aim to identify functional modules from biological networks that are enriched by top GWAS findings. Although gene functions are relevant to tissue context, most existing methods analyze tissue-free networks without reflecting phenotypic specificity. RESULTS: We propose a novel module identification framework for imaging genetic studies using the tissue-specific functional interaction network. Our method includes three steps: (i) re-prioritize imaging GWAS findings by applying machine learning methods to incorporate network topological information and enhance the connectivity among top genes; (ii) detect densely connected modules based on interactions among top re-prioritized genes; and (iii) identify phenotype-relevant modules enriched by top GWAS findings. We demonstrate our method on the GWAS of [18F]FDG-PET measures in the amygdala region using the imaging genetic data from the Alzheimer's Disease Neuroimaging Initiative, and map the GWAS results onto the amygdala-specific functional interaction network. The proposed network-based GWAS method can effectively detect densely connected modules enriched by top GWAS findings. Tissue-specific functional network can provide precise context to help explore the collective effects of genes with biologically meaningful interactions specific to the studied phenotype. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at http://www.iu.edu/shenlab/tools/gwasmodule/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xiaohui Yao, Kefei Liu 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Casey S. Greene, Jason H. Moore, Andrew J. Saykin, Li Shen 0001
Bioinform.8
2016 Exploring the coevolution of predator and prey morphology and behavior
abstract
A common idiom in biology education states, "Eyes in the front, the animal hunts. Eyes on the side, the animal hides." In this paper, we explore one possible explanation for why predators tend to have forward-facing, high-acuity visual systems. We do so using an agent-based computational model of evolution, where predators and prey interact and adapt their behavior and morphology to one another over successive generations of evolution. In this model, we observe a coevolutionary cycle between prey swarming behavior and the predator's visual system, where the predator and prey continually adapt their visual system and behavior, respectively, over evolutionary time in reaction to one another due to the well-known "predator confusion effect." Furthermore, we provide evidence that the predator visual system is what drives this coevolutionary cycle, and suggest that the cycle could be closed if the predator evolves a hybrid visual system capable of narrow, high-acuity vision for tracking prey as well as broad, coarse vision for prey discovery. Thus, the conflicting demands imposed on a predator's visual system by the predator confusion effect could have led to the evolution of complex eyes in many predators.
Christoph Adami, Jason H. Moore, Fred C. Dyer, Arend Hintze, Randal S. Olson
ALIFE2
2016 Bicliques in Graphs with Correlated Edges: From Artificial to Biological Networks
Aaron Kershenbaum, Alicia Cutillo, Christian Darabos, Keitha A. Murray, Robert Schiaffino, Jason H. Moore
EvoApplications (1)6
2016 Automating Biomedical Data Science Through Tree-Based Pipeline Optimization
Randal S. Olson, Ryan J. Urbanowicz, Peter C. Andrews, Nicole A. Lavender, La Creis Kidd, Jason H. Moore
EvoApplications (1)6
2016 Evaluation of a Tree-based Pipeline Optimization Tool for Automating Data Science
abstract
As the field of data science continues to grow, there will be an ever-increasing demand for tools that make machine learning accessible to non-experts. In this paper, we introduce the concept of tree-based pipeline optimization for automating one of the most tedious parts of machine learning--pipeline design. We implement an open source Tree-based Pipeline Optimization Tool (TPOT) in Python and demonstrate its effectiveness on a series of simulated and real-world benchmark data sets. In particular, we show that TPOT can design machine learning pipelines that provide a significant improvement over a basic machine learning analysis while requiring little to no input nor prior knowledge from the user. We also address the tendency for TPOT to design overly complex pipelines by integrating Pareto optimization, which produces compact pipelines without sacrificing classification accuracy. As such, this work represents an important step toward fully automating machine learning pipeline design.
Randal S. Olson, Nathan Bartley, Ryan J. Urbanowicz, Jason H. Moore
GECCO4
2016 Evolution of Active Categorical Image Classification via Saccadic Eye Movement
Randal S. Olson, Jason H. Moore, Christoph Adami
PPSN2
2016 Pareto Inspired Multi-objective Rule Fitness for Noise-Adaptive Rule-Based Machine Learning
Ryan J. Urbanowicz, Randal S. Olson, Jason H. Moore
PPSN3
2016 Adapting bioinformatics curricula for big data
abstract
Modern technologies are capable of generating enormous amounts of data that measure complex biological systems. Computational biologists and bioinformatics scientists are increasingly being asked to use these data to reveal key systems-level properties. We review the extent to which curricula are changing in the era of big data. We identify key competencies that scientists dealing with big data are expected to possess across fields, and we use this information to propose courses to meet these growing needs. While bioinformatics programs have traditionally trained students in data-intensive science, we identify areas of particular biological, computational and statistical emphasis important for this era that can be incorporated into existing curricula. For each area, we propose a course structured around these topics, which can be adapted in whole or in parts into existing curricula. In summary, specific challenges associated with big data provide an important opportunity to update existing curricula, but we do not foresee a wholesale redesign of bioinformatics training programs.
Anna C. Greene, Kristine A. Giffin, Casey S. Greene, Jason H. Moore
Briefings Bioinform.4
2016 Structured sparse canonical correlation analysis for brain imaging genetics: an improved GraphNet method
abstract
MOTIVATION: Structured sparse canonical correlation analysis (SCCA) models have been used to identify imaging genetic associations. These models either use group lasso or graph-guided fused lasso to conduct feature selection and feature grouping simultaneously. The group lasso based methods require prior knowledge to define the groups, which limits the capability when prior knowledge is incomplete or unavailable. The graph-guided methods overcome this drawback by using the sample correlation to define the constraint. However, they are sensitive to the sign of the sample correlation, which could introduce undesirable bias if the sign is wrongly estimated. RESULTS: We introduce a novel SCCA model with a new penalty, and develop an efficient optimization algorithm. Our method has a strong upper bound for the grouping effect for both positively and negatively correlated features. We show that our method performs better than or equally to three competing SCCA models on both synthetic and real data. In particular, our method identifies stronger canonical correlations and better canonical loading patterns, showing its promise for revealing interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/angscca/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lei Du 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001
Bioinform.7
2015 Critical properties of cellular automata with evolving network topologies
abstract
Cellular automata (CAs) in their original form are laid out on regular structures such as rings or lattices. An unsophisticated evolutionary algorithm applied to the underlying structure of the CA's connectivity is capable to significantly improve its performance solving non-trivial tasks. In this work, we study the network properties that emerge in CAs with evolving topology for the density classification problem. We compare a simple rewiring mutation operator to a more sophisticated one that allows an increase in connectivity. We also analyze the effect of initial structure in the CAs before evolution, working over the entire spectrum of regular, irregular, and random networks. We conclude that, unsurprisingly, an increase in connectivity is the driver of fitness. This also result in an increase in the clustering coefficient, and decrease in assortativity. However, our study shows that artificial evolution can also achieve high fitness in CAs with constant degree by creating shortcuts through the network, lowing the characteristic path length, and keeping the assortativity and clustering coefficient constant.
Christian Darabos, Jason H. Moore
CEC2
2015 Retooling Fitness for Noisy Problems in a Supervised Michigan-style Learning Classifier System
abstract
An accuracy-based rule fitness is a hallmark of most modern Michigan-style learning classifier systems (LCS), a powerful, flexible, and largely interpretable class of machine learners. However, rule-fitness based solely on accuracy is not ideal for identifying 'optimal' rules in supervised learning. This is particularly true for noisy problem domains where perfect rule accuracy essentially guarantees over-fitting. Rule fitness based on accuracy alone is unreliable for reflecting the global 'value' of a given rule since rule accuracy is based on a subset of the training instances. While moderate over-fitting may not dramatically hinder LCS classification or prediction performance, the interpretability of the solution is likely to suffer. Additionally, over-fitting can impede algorithm learning efficiency and leads to a larger number of rules being required to capture relationships. The present study seeks to develop an intuitive multi-objective fitness function that will encourage the discovery, preservation, and identification of 'optimal' rules through accuracy, correct coverage of training data, and the prior probability of the specified attribute states and class expressed by a given rule. We demonstrate the advantages of our proposed fitness by implementing it into the ExSTraCS algorithm and performing evaluations over a large spectrum of complex, noisy, simulated datasets.
Ryan J. Urbanowicz, Jason H. Moore
GECCO2
2015 Spectral gene set enrichment (SGSE)
abstract
BACKGROUND: Gene set testing is typically performed in a supervised context to quantify the association between groups of genes and a clinical phenotype. In many cases, however, a gene set-based interpretation of genomic data is desired in the absence of a phenotype variable. Although methods exist for unsupervised gene set testing, they predominantly compute enrichment relative to clusters of the genomic variables with performance strongly dependent on the clustering algorithm and number of clusters. RESULTS: We propose a novel method, spectral gene set enrichment (SGSE), for unsupervised competitive testing of the association between gene sets and empirical data sources. SGSE first computes the statistical association between gene sets and principal components (PCs) using our principal component gene set enrichment (PCGSE) method. The overall statistical association between each gene set and the spectral structure of the data is then computed by combining the PC-level p-values using the weighted Z-method with weights set to the PC variance scaled by Tracy-Widom test p-values. Using simulated data, we show that the SGSE algorithm can accurately recover spectral features from noisy data. To illustrate the utility of our method on real data, we demonstrate the superior performance of the SGSE method relative to standard cluster-based techniques for testing the association between MSigDB gene sets and the variance structure of microarray gene expression data. CONCLUSIONS: Unsupervised gene set testing can provide important information about the biological signal held in high-dimensional genomic data sets. Because it uses the association between gene sets and samples PCs to generate a measure of unsupervised enrichment, the SGSE method is independent of cluster or network creation algorithms and, most importantly, is able to utilize the statistical significance of PC eigenvalues to ignore elements of the data most likely to represent noise.
H. Robert Frost, Jason H. Moore
BMC Bioinform.3
2015 Delay-tolerant networks and network coding: Comparative studies on simulated and real-device experiments
Yuanzhu Peter Chen, Jiafen Liu, Walter Taylor, Jason H. Moore
Comput. Networks5
2015 An Independent Filter for Gene Set Testing Based on Spectral Enrichment
abstract
Gene set testing has become an indispensable tool for the analysis of high-dimensional genomic data. An important motivation for testing gene sets, rather than individual genomic variables, is to improve statistical power by reducing the number of tested hypotheses. Given the dramatic growth in common gene set collections, however, testing is often performed with nearly as many gene sets as underlying genomic variables. To address the challenge to statistical power posed by large gene set collections, we have developed spectral gene set filtering (SGSF), a novel technique for independent filtering of gene set collections prior to gene set testing. The SGSF method uses as a filter statistic the p-value measuring the statistical significance of the association between each gene set and the sample principal components (PCs), taking into account the significance of the associated eigenvalues. Because this filter statistic is independent of standard gene set test statistics under the null hypothesis but dependent under the alternative, the proportion of enriched gene sets is increased without impacting the type I error rate. As shown using simulated and real gene expression data, the SGSF algorithm accurately filters gene sets unrelated to the experimental outcome resulting in significantly increased gene set testing power.
H. Robert Frost, Folkert W. Asselbergs, Jason H. Moore
IEEE ACM Trans. Comput. Biol. Bioinform.4
2014 Delay-tolerant networks with network coding: How well can we simulate real devices?
abstract
Delay-tolerant networking effectively extends the network connectivity in the time domain, and endows communications devices with enhanced data transfer capabilities. Network coding on the other hand enables us to approach the information capacity of networks by allowing intermediate nodes to process data en route. Both of these were major principal breakthroughs in mobile and wireless communications in the past decade or so. As reported in this article, we are interested in how network coding battles such challenged networks as DTN from an experimental perspective. We conducted tests with both real smart mobile devices and computer simulation and found conditions where their results match. This would give us confidence of using computer simulation to study larger delay-tolerant networks with and without network coding at a much manageable cost.
Yuanzhu Peter Chen, Walter Taylor, Jason H. Moore
ICC4
2014 A Novel Structure-Aware Sparse Learning Algorithm for Brain Imaging Genetics
Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001
MICCAI (3)7
2014 Population Exploration on Genotype Networks in Genetic Programming
Ting Hu 0001, Wolfgang Banzhaf, Jason H. Moore
PPSN3
2014 An Extended Michigan-Style Learning Classifier System for Flexible Supervised Learning, Classification, and Data Mining
Ryan J. Urbanowicz, Gediminas Bertasius, Jason H. Moore
PPSN3
2014 The Effects of Recombination on Phenotypic Exploration and Robustness in Evolution
abstract
Recombination is a commonly used genetic operator in artificial and computational evolutionary systems. It has been empirically shown to be essential for evolutionary processes. However, little has been done to analyze the effects of recombination on quantitative genotypic and phenotypic properties. The majority of studies only consider mutation, mainly due to the more serious consequences of recombination in reorganizing entire genomes. Here we adopt methods from evolutionary biology to analyze a simple, yet representative, genetic programming method, linear genetic programming. We demonstrate that recombination has less disruptive effects on phenotype than mutation, that it accelerates novel phenotypic exploration, and that it particularly promotes robust phenotypes and evolves genotypic robustness and synergistic epistasis. Our results corroborate an explanation for the prevalence of recombination in complex living organisms, and helps elucidate a better understanding of the evolutionary mechanisms involved in the design of complex artificial evolutionary systems and intelligent algorithms.
Ting Hu 0001, Wolfgang Banzhaf, Jason H. Moore
Artif. Life3
2014 Robustness, Evolvability, and the Logic of Genetic Regulation
abstract
In gene regulatory circuits, the expression of individual genes is commonly modulated by a set of regulating gene products, which bind to a gene's cis-regulatory region. This region encodes an input-output function, referred to as signal-integration logic, that maps a specific combination of regulatory signals (inputs) to a particular expression state (output) of a gene. The space of all possible signal-integration functions is vast and the mapping from input to output is many-to-one: For the same set of inputs, many functions (genotypes) yield the same expression output (phenotype). Here, we exhaustively enumerate the set of signal-integration functions that yield identical gene expression patterns within a computational model of gene regulatory circuits. Our goal is to characterize the relationship between robustness and evolvability in the signal-integration space of regulatory circuits, and to understand how these properties vary between the genotypic and phenotypic scales. Among other results, we find that the distributions of genotypic robustness are skewed, so that the majority of signal-integration functions are robust to perturbation. We show that the connected set of genotypes that make up a given phenotype are constrained to specific regions of the space of all possible signal-integration functions, but that as the distance between genotypes increases, so does their capacity for unique innovations. In addition, we find that robust phenotypes are (i) evolvable, (ii) easily identified by random mutation, and (iii) mutationally biased toward other robust phenotypes. We explore the implications of these latter observations for mutation-based evolution by conducting random walks between randomly chosen source and target phenotypes. We demonstrate that the time required to identify the target phenotype is independent of the properties of the source phenotype.
Joshua L. Payne, Jason H. Moore
Artif. Life2
2014 Optimization of gene set annotations via entropy minimization over variable clusters (EMVC)
abstract
MOTIVATION: Gene set enrichment has become a critical tool for interpreting the results of high-throughput genomic experiments. Inconsistent annotation quality and lack of annotation specificity, however, limit the statistical power of enrichment methods and make it difficult to replicate enrichment results across biologically similar datasets. RESULTS: We propose a novel algorithm for optimizing gene set annotations to best match the structure of specific empirical data sources. Our proposed method, entropy minimization over variable clusters (EMVC), filters the annotations for each gene set to minimize a measure of entropy across disjoint gene clusters computed for a range of cluster sizes over multiple bootstrap resampled datasets. As shown using simulated gene sets with simulated data and Molecular Signatures Database collections with microarray gene expression data, the EMVC algorithm accurately filters annotations unrelated to the experimental outcome resulting in increased gene set enrichment power and better replication of enrichment results. AVAILABILITY AND IMPLEMENTATION: http://cran.r-project.org/web/packages/EMVC/index.html.
H. Robert Frost, Jason H. Moore
Bioinform.2
2014 Transcriptome-guided amyloid imaging genetic analysis via a novel structured sparse learning algorithm
abstract
MOTIVATION: Imaging genetics is an emerging field that studies the influence of genetic variation on brain structure and function. The major task is to examine the association between genetic markers such as single-nucleotide polymorphisms (SNPs) and quantitative traits (QTs) extracted from neuroimaging data. The complexity of these datasets has presented critical bioinformatics challenges that require new enabling tools. Sparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, most of the existing SCCA algorithms are designed using the soft thresholding method, which assumes that the input features are independent from one another. This assumption clearly does not hold for the imaging genetic data. In this article, we propose a new knowledge-guided SCCA algorithm (KG-SCCA) to overcome this limitation as well as improve learning results by incorporating valuable prior knowledge. RESULTS: The proposed KG-SCCA method is able to model two types of prior knowledge: one as a group structure (e.g. linkage disequilibrium blocks among SNPs) and the other as a network structure (e.g. gene co-expression network among brain regions). The new model incorporates these prior structures by introducing new regularization terms to encourage weight similarity between grouped or connected features. A new algorithm is designed to solve the KG-SCCA model without imposing the independence constraint on the input features. We demonstrate the effectiveness of our algorithm with both synthetic and real data. For real data, using an Alzheimer's disease (AD) cohort, we examine the imaging genetic associations between all SNPs in the APOE gene (i.e. top AD gene) and amyloid deposition measures among cortical regions (i.e. a major AD hallmark). In comparison with a widely used SCCA implementation, our KG-SCCA algorithm produces not only improved cross-validation performances but also biologically meaningful results. AVAILABILITY: Software is freely available on request.
Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Jason H. Moore, Andrew J. Saykin, Li Shen 0001
Bioinform.6
2014 Phenotypic Robustness and the Assortativity Signature of Human Transcription Factor Networks
abstract
Many developmental, physiological, and behavioral processes depend on the precise expression of genes in space and time. Such spatiotemporal gene expression phenotypes arise from the binding of sequence-specific transcription factors (TFs) to DNA, and from the regulation of nearby genes that such binding causes. These nearby genes may themselves encode TFs, giving rise to a transcription factor network (TFN), wherein nodes represent TFs and directed edges denote regulatory interactions between TFs. Computational studies have linked several topological properties of TFNs - such as their degree distribution - with the robustness of a TFN's gene expression phenotype to genetic and environmental perturbation. Another important topological property is assortativity, which measures the tendency of nodes with similar numbers of edges to connect. In directed networks, assortativity comprises four distinct components that collectively form an assortativity signature. We know very little about how a TFN's assortativity signature affects the robustness of its gene expression phenotype to perturbation. While recent theoretical results suggest that increasing one specific component of a TFN's assortativity signature leads to increased phenotypic robustness, the biological context of this finding is currently limited because the assortativity signatures of real-world TFNs have not been characterized. It is therefore unclear whether these earlier theoretical findings are biologically relevant. Moreover, it is not known how the other three components of the assortativity signature contribute to the phenotypic robustness of TFNs. Here, we use publicly available DNaseI-seq data to measure the assortativity signatures of genome-wide TFNs in 41 distinct human cell and tissue types. We find that all TFNs share a common assortativity signature and that this signature confers phenotypic robustness to model TFNs. Lastly, we determine the extent to which each of the four components of the assortativity signature contributes to this robustness.
Dov A. Pechenick, Joshua L. Payne, Jason H. Moore
PLoS Comput. Biol.3
2013 Robustness and Evolvability of Recombination in Linear Genetic Programming
Ting Hu 0001, Wolfgang Banzhaf, Jason H. Moore
EuroGP3
2013 Research and applications: An information-gain approach to detecting three-way epistatic interactions in genetic association studies
abstract
BACKGROUND: Epistasis has been historically used to describe the phenomenon that the effect of a given gene on a phenotype can be dependent on one or more other genes, and is an essential element for understanding the association between genetic and phenotypic variations. Quantifying epistasis of orders higher than two is very challenging due to both the computational complexity of enumerating all possible combinations in genome-wide data and the lack of efficient and effective methodologies. OBJECTIVES: In this study, we propose a fast, non-parametric, and model-free measure for three-way epistasis. METHODS: Such a measure is based on information gain, and is able to separate all lower order effects from pure three-way epistasis. RESULTS: Our method was verified on synthetic data and applied to real data from a candidate-gene study of tuberculosis in a West African population. In the tuberculosis data, we found a statistically significant pure three-way epistatic interaction effect that was stronger than any lower-order associations. CONCLUSION: Our study provides a methodological basis for detecting and characterizing high-order gene-gene interactions in genetic association studies.
Ting Hu 0001, Yuanzhu Peter Chen, Jeff Kiralis, Ryan L. Collins, Christian Wejse, Giorgio Sirugo, Scott M. Williams, Jason H. Moore
J. Am. Medical Informatics Assoc.8
2013 Research and applications: Role of genetic heterogeneity and epistasis in bladder cancer susceptibility and outcome: a learning classifier system approach
abstract
BACKGROUND AND OBJECTIVE: Detecting complex patterns of association between genetic or environmental risk factors and disease risk has become an important target for epidemiological research. In particular, strategies that provide multifactor interactions or heterogeneous patterns of association can offer new insights into association studies for which traditional analytic tools have had limited success. MATERIALS AND METHODS: To concurrently examine these phenomena, previous work has successfully considered the application of learning classifier systems (LCSs), a flexible class of evolutionary algorithms that distributes learned associations over a population of rules. Subsequent work dealt with the inherent problems of knowledge discovery and interpretation within these algorithms, allowing for the characterization of heterogeneous patterns of association. Whereas these previous advancements were evaluated using complex simulation studies, this study applied these collective works to a 'real-world' genetic epidemiology study of bladder cancer susceptibility. RESULTS AND DISCUSSION: We replicated the identification of previously characterized factors that modify bladder cancer risk--namely, single nucleotide polymorphisms from a DNA repair gene, and smoking. Furthermore, we identified potentially heterogeneous groups of subjects characterized by distinct patterns of association. Cox proportional hazard models comparing clinical outcome variables between the cases of the two largest groups yielded a significant, meaningful difference in survival time in years (survivorship). A marginally significant difference in recurrence time was also noted. These results support the hypothesis that an LCS approach can offer greater insight into complex patterns of association. CONCLUSIONS: This methodology appears to be well suited to the dissection of disease heterogeneity, a key component in the advancement of personalized medicine.
Ryan J. Urbanowicz, Angeline S. Andrew, Margaret R. Karagas, Jason H. Moore
J. Am. Medical Informatics Assoc.4
2013 Complex and dynamic population structures: synthesis, open questions, and future directions
Joshua L. Payne, Mario Giacobini, Jason H. Moore
Soft Comput.3
2012 Instance-linked attribute tracking and feedback for michigan-style supervised learning classifier systems
abstract
The application of learning classifier systems (LCSs) to classification and data mining in genetic association studies has been the target of previous work. Recent efforts have focused on: (1) correctly discriminating between predictive and non-predictive attributes, and (2) detecting and characterizing epistasis (attribute interaction) and heterogeneity. While the solutions evolved by Michigan-style LCSs (M-LCSs) are conceptually well suited to address these phenomena, the explicit characterization of heterogeneity remains a particular challenge. In this study we introduce attribute tracking, a mechanism akin to memory, for supervised learning in M-LCSs. Given a finite training set, a vector of accuracy scores is maintained for each instance in the data. Post-training, we apply these scores to characterize patterns of association in the dataset. Additionally we introduce attribute feedback to the mutation and crossover mechanisms, probabilistically directing rule generalization based on an instance's tracking scores. We find that attribute tracking combined with clustering and visualization facilitates the characterization of epistasis and heterogeneity while uniquely linking individual instances in the dataset to etiologically heterogeneous subgroups. Moreover, these analyses demonstrate that attribute feedback significantly improves test accuracy, efficient generalization, run time, and the power to discriminate between predictive and non-predictive attributes in the presence of heterogeneity.
Ryan J. Urbanowicz, Ambrose Granizo-Mackenzie, Jason H. Moore
GECCO3
2012 Using Expert Knowledge to Guide Covering and Mutation in a Michigan Style Learning Classifier System to Detect Epistasis and Heterogeneity
Ryan J. Urbanowicz, Delaney Granizo-MacKenzie, Jason H. Moore
PPSN (1)3
2012 Measuring the microbiome: perspectives on advances in DNA-based techniques for exploring microbial life
abstract
This article reviews recent advances in 'microbiome studies': molecular, statistical and graphical techniques to explore and quantify how microbial organisms affect our environments and ourselves given recent increases in sequencing technology. Microbiome studies are moving beyond mere inventories of specific ecosystems to quantifications of community diversity and descriptions of their ecological function. We review the last 24 months of progress in this sort of research, and anticipate where the next 2 years will take us. We hope that bioinformaticians will find this a helpful springboard for new collaborations with microbiologists.
James A. Foster, John Bunge, Jack A. Gilbert, Jason H. Moore
Briefings Bioinform.4
2012 Chapter 11: Genome-Wide Association Studies
abstract
Genome-wide association studies (GWAS) have evolved over the last ten years into a powerful tool for investigating the genetic architecture of human disease. In this work, we review the key concepts underlying GWAS, including the architecture of common diseases, the structure of common human genetic variation, technologies for capturing genetic information, study designs, and the statistical methods used for data analysis. We also look forward to the future beyond GWAS.
William S. Bush, Jason H. Moore
PLoS Comput. Biol.2
2011 Robustness, Evolvability, and Accessibility in Linear Genetic Programming
Ting Hu 0001, Joshua L. Payne, Wolfgang Banzhaf, Jason H. Moore
EuroGP4
2011 Characterizing Genetic Interactions in Human Disease Association Studies Using Statistical Epistasis Networks
abstract
BACKGROUND: Epistasis is recognized ubiquitous in the genetic architecture of complex traits such as disease susceptibility. Experimental studies in model organisms have revealed extensive evidence of biological interactions among genes. Meanwhile, statistical and computational studies in human populations have suggested non-additive effects of genetic variation on complex traits. Although these studies form a baseline for understanding the genetic architecture of complex traits, to date they have only considered interactions among a small number of genetic variants. Our goal here is to use network science to determine the extent to which non-additive interactions exist beyond small subsets of genetic variants. We infer statistical epistasis networks to characterize the global space of pairwise interactions among approximately 1500 Single Nucleotide Polymorphisms (SNPs) spanning nearly 500 cancer susceptibility genes in a large population-based study of bladder cancer. RESULTS: The statistical epistasis network was built by linking pairs of SNPs if their pairwise interactions were stronger than a systematically derived threshold. Its topology clearly differentiated this real-data network from networks obtained from permutations of the same data under the null hypothesis that no association exists between genotype and phenotype. The network had a significantly higher number of hub SNPs and, interestingly, these hub SNPs were not necessarily with high main effects. The network had a largest connected component of 39 SNPs that was absent in any other permuted-data networks. In addition, the vertex degrees of this network were distinctively found following an approximate power-law distribution and its topology appeared scale-free. CONCLUSIONS: In contrast to many existing techniques focusing on high main-effect SNPs or models of several interacting SNPs, our network approach characterized a global picture of gene-gene interactions in a population-based genetic data. The network was built using pairwise interactions, and its distinctive network topology and large connected components indicated joint effects in a large set of SNPs. Our observations suggested that this particular statistical epistasis network captured important features of the genetic architecture of bladder cancer that have not been described previously.
Ting Hu 0001, Nicholas A. Sinnott-Armstrong, Jeff Kiralis, Angeline S. Andrew, Margaret R. Karagas, Jason H. Moore
BMC Bioinform.6
2010 Sexual Recombination in Self-Organizing Interaction Networks
Joshua L. Payne, Jason H. Moore
EvoApplications (1)2
2010 Fast genome-wide epistasis analysis using ant colony optimization for multifactor dimensionality reduction analysis on graphics processing units
abstract
Epistasis, or non-linear gene-to-gene interaction, is now thought to be at the heart of many common human diseases. A popular algorithm to detect epistasis is Multifactor Dimensionality Reduction (MDR), which exhaustively searches to determine an optimal classification. This exhaustive search is combinatorial in complexity and does not scale efficiently to large datasets. Ant Colony Opimization (ACO) is a technique to reduce this complexity by exploiting expert knowledge to spend more time looking at most likely candidates for the optimal classification. Graphics Processing Units (GPUs) are highly-parallel integrated circuits able to execute arbitrary code. The authors implemented ACO MDR on GPUs and compared it to both a Java ACO implementation and an exhaustive C++ implementation. The performance advantage of GPUs, combined with the added computational efficiency of a heuristic evolutionary algorithm such as ACO, allow larger scale problems to be tackled, something that is becoming critical with the advances in high throughput genome sequencing.
Nicholas A. Sinnott-Armstrong, Casey S. Greene, Jason H. Moore
GECCO3
2010 The application of michigan-style learning classifiersystems to address genetic heterogeneity and epistasisin association studies
abstract
Genetic epidemiologists, tasked with the disentanglement of genotype-to-phenotype mappings, continue to struggle with a variety of phenomena which obscure the underlying etiologies of common complex diseases. For genetic association studies, genetic heterogeneity (GH) and epistasis (gene-gene interactions) epitomize well recognized phenomenon which represent a difficult, but accessible challenge for computational biologists. While progress has been made addressing epistasis, methods for dealing with GH tend to "side-step" the problem, limited by a dependence on potentially arbitrary cutoffs/covariates, and a loss in power synonymous with data stratification. In the present study, we explore an alternative strategy (Learning Classifier Systems (LCSs)) as a direct approach for the characterization, and modeling of disease in the presence of both GH and epistasis. This evaluation involves (1) implementing standardized versions of existing Michigan-Style LCSs (XCS, MCS, and UCS), (2) examining major run parameters, and (3) performing quantitative and qualitative evaluations across a spectrum of simulated datasets. The results of this study highlight the strengths and weaknesses of the Michigan LCS architectures examined, providing proof of principle for the application of LCSs to the GH/epistasis problem, and laying the foundation for the development of an LCS algorithm specifically designed to address GH.
Ryan J. Urbanowicz, Jason H. Moore
GECCO2
2010 The Application of Pittsburgh-Style Learning Classifier Systems to Address Genetic Heterogeneity and Epistasis in Association Studies
Ryan J. Urbanowicz, Jason H. Moore
PPSN (1)2
2010 Multifactor dimensionality reduction for graphics processing units enables genome-wide testing of epistasis in sporadic ALS
abstract
MOTIVATION: Epistasis, the presence of gene-gene interactions, has been hypothesized to be at the root of many common human diseases, but current genome-wide association studies largely ignore its role. Multifactor dimensionality reduction (MDR) is a powerful model-free method for detecting epistatic relationships between genes, but computational costs have made its application to genome-wide data difficult. Graphics processing units (GPUs), the hardware responsible for rendering computer games, are powerful parallel processors. Using GPUs to run MDR on a genome-wide dataset allows for statistically rigorous testing of epistasis. RESULTS: The implementation of MDR for GPUs (MDRGPU) includes core features of the widely used Java software package, MDR. This GPU implementation allows for large-scale analysis of epistasis at a dramatically lower cost than the standard CPU-based implementations. As a proof-of-concept, we applied this software to a genome-wide study of sporadic amyotrophic lateral sclerosis (ALS). We discovered a statistically significant two-SNP classifier and subsequently replicated the significance of these two SNPs in an independent study of ALS. MDRGPU makes the large-scale analysis of epistasis tractable and opens the door to statistically rigorous testing of interactions in genome-wide datasets. AVAILABILITY: MDRGPU is open source and available free of charge from http://www.sourceforge.net/projects/mdr.
Casey S. Greene, Nicholas A. Sinnott-Armstrong, Daniel S. Himmelstein, Paul J. Park, Jason H. Moore, Brent T. Harris
Bioinform.5
2010 Bioinformatics challenges for genome-wide association studies
abstract
MOTIVATION: The sequencing of the human genome has made it possible to identify an informative set of >1 million single nucleotide polymorphisms (SNPs) across the genome that can be used to carry out genome-wide association studies (GWASs). The availability of massive amounts of GWAS data has necessitated the development of new biostatistical methods for quality control, imputation and analysis issues including multiple testing. This work has been successful and has enabled the discovery of new associations that have been replicated in multiple studies. However, it is now recognized that most SNPs discovered via GWAS have small effects on disease susceptibility and thus may not be suitable for improving health care through genetic testing. One likely explanation for the mixed results of GWAS is that the current biostatistical analysis paradigm is by design agnostic or unbiased in that it ignores all prior knowledge about disease pathobiology. Further, the linear modeling framework that is employed in GWAS often considers only one SNP at a time thus ignoring their genomic and environmental context. There is now a shift away from the biostatistical approach toward a more holistic approach that recognizes the complexity of the genotype-phenotype relationship that is characterized by significant heterogeneity and gene-gene and gene-environment interaction. We argue here that bioinformatics has an important role to play in addressing the complexity of the underlying genetic basis of common human diseases. The goal of this review is to identify and discuss those GWAS challenges that will require computational methods.
Jason H. Moore, Folkert W. Asselbergs, Scott M. Williams
Bioinform.1
2009 Nature-inspired algorithms for the genetic analysis of epistasis in common human diseases: Theoretical assessment of wrapper vs. filter approaches
abstract
In human genetics, new technological methods allow researchers to collect a wealth of information about genetic variation among individuals quickly and relatively inexpensively. Studies examining more than one half of a million points of genetic variation are the new standard. Quickly analyzing these data to discover single gene effects is both feasible and often done. Unfortunately as our understanding of common human disease grows, we now believe it is likely that an individual's risk of these common diseases is not determined by simple single gene effects. Instead it seems likely that risk will be determined by nonlinear gene-gene interactions, also known as epistasis. Unfortunately searching for these nonlinear effects requires either effective search strategies or exhaustive search. Previously we have employed both filter and nature-inspired probabilistic search wrapper approaches such as genetic programming (GP) and ant colony optimization (ACO) to this problem. We have discovered that for this problem, expert knowledge is critical if we are to discover these interactions. Here we theoretically analyze both an expert knowledge filter and a simple expert-knowledge-aware wrapper. We show that under certain assumptions, the filter strategy leads to the highest power. Finally we discuss the implications of this work for this type of problem, and discuss how probabilistic search strategies which outperform a filtering approach may be designed.
Casey S. Greene, Jeff Kiralis, Jason H. Moore
IEEE Congress on Evolutionary Computation3
2009 Sensible initialization using expert knowledge for genome-wide analysis of epistasis using genetic programming
abstract
For biomedical researchers it is now possible to measure large numbers of DNA sequence variations across the human genome. Measuring hundreds of thousands of variations is now routine, but single variations which consistently predict an individual's risk of common human disease have proven elusive. Instead of single variants determining the risk of common human diseases, it seems more likely that disease risk is best modeled by interactions between biological components. The evolutionary computing challenge now is to effectively explore interactions in these large datasets and identify combinations of variations which are robust predictors of common human diseases such as bladder cancer. One promising approach to this problem is genetic programming (GP). A GP approach for this problem will use darwinian inspired evolution to evolve programs which find and model attribute interactions which predict an individual's risk of common human diseases. The goal of this study is to develop and evaluate two initializers for this domain. We develop a probabilistic initializer which uses expert knowledge to select attributes and an enumerative initializer which maximizes attribute diversity in the generated population.We compare these initializers to a random initializer which displays no preference for attributes. We show that the expert-knowledge-aware probabilistic initializer significantly outperforms both the random initializer and the enumerative initializer.We discuss implications of these results for the design of GP strategies which are able to detect and characterize predictors of common human diseases.
Casey S. Greene, Bill C. White, Jason H. Moore
IEEE Congress on Evolutionary Computation3
2009 Development and evaluation of an open-ended computational evolution system for the creation of digital organisms with complex genetic architecture
abstract
Epistasis, or gene-gene interaction, is a ubiquitous phenomenon that is inadequately addressed in human genetic studies. There are few tools that can accurately identify high-order epistatic interactions, and there is a lack of general understanding as to how epistatic interactions fit into genetic architecture. Here we approach both problems through the lens of genetic programming (GP). It has recently been proposed that increasing open-endedness of GP will result in more complex solutions that better acknowledge the complexity of human genetic datasets. Moreover, the solutions evolved in open-ended GP can serve as model organisms in which to study general effects of epistasis on phenotype. Here we introduce a prototype computational evolution system that implements an open-ended GP and generates organisms that display epistatic interactions. These interactions are significantly more prevalent and have a greater effect on fitness than epistatic interactions in organisms generated in the absence of selection.
Anna L. Tyler, Bill C. White, Casey S. Greene, Paul C. Andrews, Richard Cowper-Sallari, Jason H. Moore
IEEE Congress on Evolutionary Computation6
2009 Environmental noise improves epistasis models of genetic data discovered using a computational evolution system
abstract
Common human diseases likely result from nonlinear interactions between multiple DNA sequence variations. One goal of human genetics is to use data mining and machine learning methods to identify combinations of genetic variations that are predictive of discrete measures of health in human population data. "Artificial evolution" approaches loosely based on real biological processes have been developed and applied in this domain, but it has recently been suggested that "computational evolution" approaches which incorporate additional biological and evolutionary complexity into existing algorithms will be more likely to solve problems of interest to biologists and biomedical researchers. Here we introduce a method to evolve compact solutions by adding environmental noise to a dataset during fitness evaluation. In ecological systems a highly specialized organism can fail to thrive as the environment changes. By introducing numerous small changes into training data, i.e. the environment, during evolution we similarly drive selection towards more general solutions. We show that this improves the power of the computational evolution system when modest amounts of noise are used. Furthermore, this method of changing the environment in which fitness is evaluated with small perturbations fits within the computational evolution framework and is an effective method of controlling solution size for problems where the data are likely to be noisy.
Casey S. Greene, Douglas P. Hill, Jason H. Moore
GECCO3
2008 Using expert knowledge in initialization for genome-wide analysis of epistasis using genetic programming
abstract
In human genetics it is now possible to measure large numbers of DNA sequence variations across the human genome. Given current knowledge about biological networks and disease processes it seems likely that disease risk can best be modeled by interactions between biological components, which may be examined as interacting DNA sequence variations. The machine learning challenge is to e.ectively explore interactions in these datasets to identify combinations of variations which are predictive of common human diseases. Genetic programming is a promising approach to this problem. The goal of this study is to examine the role that an expert knowledge aware initializer can play in the framework of genetic programming. We show that this expert knowledge aware initializer outperforms both a random initializer and an enumerative initializer.
Casey S. Greene, Bill C. White, Jason H. Moore
GECCO3
2008 Mask functions for the symbolic modeling of epistasis using genetic programming
abstract
The study of common, complex multifactorial diseases in genetic epidemiology is complicated by nonlinearity in the genotype-to-phenotype mapping relationship that is due, in part, to epistasis or gene-gene interactions. Symobolic discriminant analysis (SDA) is a flexible modeling approach which uses genetic programming (GP) to evolve an optimal predictive model using a predefined collection of mathematical functions, constants, and attributes. This has been shown to be an effective strategy for modeling epistasis. In the present study, we introduce the genetic "mask" as a novel building block which exploits expert knowledge in the form of a pre-constructed relationship between two attributes. The goal of this study was to determine whether the availability of "mask" building blocks improves SDA performance. The results of this study support the idea that pre-processing data improves GP performance.
Ryan J. Urbanowicz, Nate Barney, Bill C. White, Jason H. Moore
GECCO4
2007 Towards human-human-computer interaction for biologically-inspired problem-solving in human genetics
abstract
Genetic programming (GP) shows great promise for solving complex problems in human genetics. Unfortunately, many of these methods are not accessible to biologists. This is partly due to the complexity of the algorithms that limit their ready adoption and integration into an analysis or modeling paradigm that might otherwise only use univariate statistical methods.allThis is also partly due to the lack of user-friendly, open-source, platform-independent, and freely-available software packages that are designed to be used by biologists for routine analysis. It is our objective to develop, distribute and support a comprehensive software package that puts powerful GP methods for genetic analysis in the hands of geneticists. It is our working hypothesis that the most effective use of such a software package would result from interactive analysis by both a biologist and a computer scientist (i.e. human-human-computer interaction).allWe summarize briefly here the design and implementation of an open-source software package called Symbolic Modeler (SyMod) that seeks to facilitate geneticist-bioinformaticist-computer interactions for problem solving in human genetics. More information can be found at www.epistasis.org or www.symbolicmodeler.org.
Jason H. Moore, Nate Barney, Bill C. White
GECCO1
2007 Evaporative cooling feature selection for genotypic data involving interactions
abstract
MOTIVATION: The development of genome-wide capabilities for genotyping has led to the practical problem of identifying the minimum subset of genetic variants relevant to the classification of a phenotype. This challenge is especially difficult in the presence of attribute interactions, noise and small sample size. METHODS: Analogous to the physical mechanism of evaporation, we introduce an evaporative cooling (EC) feature selection algorithm that seeks to obtain a subset of attributes with the optimum information temperature (i.e. the least noise). EC uses an attribute quality measure analogous to thermodynamic free energy that combines Relief-F and mutual information to evaporate (i.e. remove) noise features, leaving behind a subset of attributes that contain DNA sequence variations associated with a given phenotype. RESULTS: EC is able to identify functional sequence variations that involve interactions (epistasis) between other sequence variations that influence their association with the phenotype. This ability is demonstrated on simulated genotypic data with attribute interactions and on real genotypic data from individuals who experienced adverse events following smallpox vaccination. The EC formalism allows us to combine information entropy, energy and temperature into a single information free energy attribute quality measure that balances interaction and main effects. AVAILABILITY: Open source software, written in Java, is freely available upon request.
Brett A. McKinney, David M. Reif, Bill C. White, James E. Crowe Jr., Jason H. Moore
Bioinform.5
2006 Feature Selection using a Random Forests Classifier for the Integrated Analysis of Multiple Data Types
abstract
Complex clinical phenotypes arise from the concerted interactions among the myriad components of a biological system. Therefore, comprehensive models can only be developed through the integrated study of multiple types of experimental data gathered from the system in question. The Random Foreststrade(RF) method is adept at identifying relevant features having only slight main effects in high-dimensional data. This method is well-suited to integrated analysis, as relevant attributes may be selected from categorical or continuous data, and there may be interactions across data types. RF is a natural approach for studying gene-gene, gene-protein, or protein-protein interactions because importance scores for particular attributes take interactions into account. Thus, Random Forests is a promising solution to the analysis challenge posed by high-dimensional datasets including interactions among attributes of different types. In this study, we characterize the performance of RF on a range of simulated genetic and/or proteomic datasets. We compare the performance of RF in identifying relevant attributes when given genetic data alone, proteomic data alone, or a combined dataset of genetic plus proteomic data. Our results indicate that utilizing multiple data types is beneficial when the disease model is complex and the phenotypic outcome-associated data type is unknown. The results of this study also show that RF is adept at identifying relevant features in high-dimensional data with small main effects and low heritability
David M. Reif, Alison A. Motsinger-Reif, Brett A. McKinney, James E. Crowe Jr., Jason H. Moore
CIBCB5
2006 Exploiting Expert Knowledge in Genetic Programming for Genome-Wide Genetic Analysis
Jason H. Moore, Bill C. White
PPSN1
2006 Dissecting trait heterogeneity: a comparison of three clustering methods applied to genotypic data
abstract
BACKGROUND: Trait heterogeneity, which exists when a trait has been defined with insufficient specificity such that it is actually two or more distinct traits, has been implicated as a confounding factor in traditional statistical genetics of complex human disease. In the absence of detailed phenotypic data collected consistently in combination with genetic data, unsupervised computational methodologies offer the potential for discovering underlying trait heterogeneity. The performance of three such methods--Bayesian Classification, Hypergraph-Based Clustering, and Fuzzy k-Modes Clustering--appropriate for categorical data were compared. Also tested was the ability of these methods to detect trait heterogeneity in the presence of locus heterogeneity and/or gene-gene interaction, which are two other complicating factors in discovering genetic models of complex human disease. To determine the efficacy of applying the Bayesian Classification method to real data, the reliability of its internal clustering metrics at finding good clusterings was evaluated using permutation testing. RESULTS: Bayesian Classification outperformed the other two methods, with the exception that the Fuzzy k-Modes Clustering performed best on the most complex genetic model. Bayesian Classification achieved excellent recovery for 75% of the datasets simulated under the simplest genetic model, while it achieved moderate recovery for 56% of datasets with a sample size of 500 or more (across all simulated models) and for 86% of datasets with 10 or fewer nonfunctional loci (across all simulated models). Neither Hypergraph Clustering nor Fuzzy k-Modes Clustering achieved good or excellent cluster recovery for a majority of datasets even under a restricted set of conditions. When using the average log of class strength as the internal clustering metric, the false positive rate was controlled very well, at three percent or less for all three significance levels (0.01, 0.05, 0.10), and the false negative rate was acceptably low (18 percent) for the least stringent significance level of 0.10. CONCLUSION: Bayesian Classification shows promise as an unsupervised computational method for dissecting trait heterogeneity in genotypic data. Its control of false positive and false negative rates lends confidence to the validity of its results. Further investigation of how different parameter settings may improve the performance of Bayesian Classification, especially under more complex genetic models, is ongoing.
Tricia A. Thornton-Wells, Jason H. Moore, Jonathan L. Haines
BMC Bioinform.2
2005 A statistical comparison of grammatical evolution strategies in the domain of human genetics
abstract
Detecting and characterizing genetic predictors of human disease susceptibility is an important goal in human genetics. New chip-based technologies are available that facilitate the measurement of thousands of DNA sequence variations across the human genome. Biologically-inspired stochastic search algorithms are expected to play an important role in the analysis of these high-dimensional datasets. We simulated datasets with up to 6000 attributes using two different genetic models and statistically compared the performance of grammatical evolution, grammatical swarm, and random search for building symbolic discriminant functions. We found no statistical difference among search algorithms within this specific domain.
Bill C. White, David M. Reif, Joshua C. Gilbert, Jason H. Moore
Congress on Evolutionary Computation4
2005 A statistical comparison of grammatical evolution strategies in the domain of human genetics
abstract
Detecting and characterizing genetic predictors of human disease susceptibility is an important goal in human genetics. New chip-based technologies are available that facilitate the measurement of thousands of DNA sequence variations across the human genome. Biologically-inspired stochastic search algorithms are expected to play an important role in the analysis of these high-dimensional datasets. We simulated datasets with up to 6000 attributes using two different genetic models and statistically compared the performance of grammatical evolution, grammatical swarm, and random search for building symbolic discriminant functions. We found no statistical difference among search algorithms within this specific domain.
Bill C. White, David M. Reif, Joshua C. Gilbert, Jason H. Moore
Congress on Evolutionary Computation4
2004 Systems Biology Modeling in Human Genetics Using Petri Nets and Grammatical Evolution
Jason H. Moore, Lance W. Hahn
GECCO (1)1
2004 Genetic Programming Neural Networks as a Bioinformatics Tool for Human Genetics
Marylyn D. Ritchie, Christopher S. Coffey, Jason H. Moore
GECCO (1)3
2004 An application of conditional logistic regression and multifactor dimensionality reduction for detecting gene-gene Interactions on risk of myocardial infarction: The importance of model validation
abstract
BACKGROUND: To examine interactions among the angiotensin converting enzyme (ACE) insertion/deletion, plasminogen activator inhibitor-1 (PAI-1) 4G/5G, and tissue plasminogen activator (t-PA) insertion/deletion gene polymorphisms on risk of myocardial infarction using data from 343 matched case-control pairs from the Physicians Health Study. We examined the data using both conditional logistic regression and the multifactor dimensionality reduction (MDR) method. One advantage of the MDR method is that it provides an internal prediction error for validation. We summarize our use of this internal prediction error for model validation. RESULTS: The overall results for the two methods were consistent, with both suggesting an interaction between the ACE I/D and PAI-1 4G/5G polymorphisms. However, using ten-fold cross validation, the 46% prediction error for the final MDR model was not significantly lower than that expected by chance. CONCLUSIONS: The significant interaction initially observed does not validate and may represent a type I error. As data-driven analytic methods continue to be developed and used to examine complex genetic interactions, it will become increasingly important to stress model validation in order to ensure that significant effects represent true relationships rather than chance findings.
Christopher S. Coffey, Patricia R. Hebert, Marylyn D. Ritchie, Harlan M. Krumholz, John Michael Gaziano, Paul M. Ridker, Nancy J. Brown, Douglas E. Vaughan, Jason H. Moore
BMC Bioinform.9
2003 Grammatical Evolution for the Discovery of Petri Net Models of Complex Genetic Systems
Jason H. Moore, Lance W. Hahn
GECCO1
2003 Complex Function Sets Improve Symbolic Discriminant Analysis of Microarray Data
David M. Reif, Bill C. White, Nancy Olsen, Thomas M. Aune, Jason H. Moore
GECCO5
2003 Multifactor dimensionality reduction software for detecting gene-gene and gene-environment interactions
abstract
MOTIVATION: Polymorphisms in human genes are being described in remarkable numbers. Determining which polymorphisms and which environmental factors are associated with common, complex diseases has become a daunting task. This is partly because the effect of any single genetic variation will likely be dependent on other genetic variations (gene-gene interaction or epistasis) and environmental factors (gene-environment interaction). Detecting and characterizing interactions among multiple factors is both a statistical and a computational challenge. To address this problem, we have developed a multifactor dimensionality reduction (MDR) method for collapsing high-dimensional genetic data into a single dimension thus permitting interactions to be detected in relatively small sample sizes. In this paper, we describe the MDR approach and an MDR software package. RESULTS: We developed a program that integrates MDR with a cross-validation strategy for estimating the classification and prediction error of multifactor models. The software can be used to analyze interactions among 2-15 genetic and/or environmental factors. The dataset may contain up to 500 total variables and a maximum of 4000 study subjects. AVAILABILITY: Information on obtaining the executable code, example data, example analysis, and documentation is available upon request. SUPPLEMENTARY INFORMATION: All supplementary information can be found at http://phg.mc.vanderbilt.edu/Software/MDR.
Lance W. Hahn, Marylyn D. Ritchie, Jason H. Moore
Bioinform.3
2003 Optimizationof neural network architecture using genetic programming improvesdetection and modeling of gene-gene interactions in studies of humandiseases
abstract
BACKGROUND: Appropriate definition of neural network architecture prior to data analysis is crucial for successful data mining. This can be challenging when the underlying model of the data is unknown. The goal of this study was to determine whether optimizing neural network architecture using genetic programming as a machine learning strategy would improve the ability of neural networks to model and detect nonlinear interactions among genes in studies of common human diseases. RESULTS: Using simulated data, we show that a genetic programming optimized neural network approach is able to model gene-gene interactions as well as a traditional back propagation neural network. Furthermore, the genetic programming optimized neural network is better than the traditional back propagation neural network approach in terms of predictive ability and power to detect gene-gene interactions when non-functional polymorphisms are present. CONCLUSION: This study suggests that a machine learning strategy for optimizing neural network architecture may be preferable to traditional trial-and-error approaches for the identification and characterization of gene-gene interactions in common, complex human diseases.
Marylyn D. Ritchie, Bill C. White, Joel S. Parker, Lance W. Hahn, Jason H. Moore
BMC Bioinform.5
2002 Application Of Genetic Algorithms To The Discovery Of Complex Models For Simulation Studies In Human Genetics
Jason H. Moore, Lance W. Hahn, Marylyn D. Ritchie, Tricia A. Thornton-Wells, Bill C. White
GECCO1
2002 Cellular Automata and Genetic Algorithms for Parallel Problem Solving in Human Genetics
Jason H. Moore, Lance W. Hahn
PPSN1
2001 Symbolic Discriminant Analysis for Mining Gene Expression Patterns
Jason H. Moore, Joel S. Parker, Lance W. Hahn
ECML1