Atul J. Butte

dblp:96/1665 · DBLP profile ↗
← Back
65ranked-venue papers
11as first author
11since 2021 · last 2024
0000-0002-7433-2740ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 63 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 A comparative study of large language model-based zero-shot inference and task-specific supervised classification of breast cancer pathology reports
abstract
OBJECTIVE: Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs could reduce the need for large-scale data annotations. MATERIALS AND METHODS: We curated a dataset of 769 breast cancer pathology reports, manually labeled with 12 categories, to compare zero-shot classification capability of the following LLMs: GPT-4, GPT-3.5, Starling, and ClinicalCamel, with task-specific supervised classification performance of 3 models: random forests, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. RESULTS: Across all 12 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, LSTM-Att (average macro F1-score of 0.86 vs 0.75), with advantage on tasks with high label imbalance. Other LLMs demonstrated poor performance. Frequent GPT-4 error categories included incorrect inferences from multiple samples and from history, and complex task design, and several LSTM-Att errors were related to poor generalization to the test set. DISCUSSION: On tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of data labeling. However, if the use of LLMs is prohibitive, the use of simpler models with large annotated datasets can provide comparable results. CONCLUSIONS: GPT-4 demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for large annotated datasets. This may increase the utilization of NLP-based variables and outcomes in clinical studies.
Madhumita Sushil, Travis Zack, Divneet Mandair, Zhiwei Zheng, Ahmed Wali, Yan-Ning Yu, Yuwei Quan, Dmytro Lituiev, Atul J. Butte
J. Am. Medical Informatics Assoc.9
2023 Aligning Synthetic Medical Images with Clinical Knowledge using Human Feedback
abstract
Generative models capable of precisely capturing nuanced clinical features in medical images hold great promise for facilitating clinical data sharing, enhancing rare disease datasets, and efficiently synthesizing (annotated) medical images at scale. Despite their potential, assessing the quality of synthetic medical images remains a challenge. While modern generative models can synthesize visually-realistic medical images, the clinical plausibility of these images may be called into question. Domain-agnostic scores, such as FID score, precision, and recall, cannot incorporate clinical knowledge and are, therefore, not suitable for assessing clinical sensibility. Additionally, there are numerous unpredictable ways in which generative models may fail to synthesize clinically plausible images, making it challenging to anticipate potential failures and design automated scores for their detection. To address these challenges, this paper introduces a pathologist-in-the-loop framework for generating clinically-plausible synthetic medical images. Our framework comprises three steps: (1) pretraining a conditional diffusion model to generate medical images conditioned on a clinical concept, (2) expert pathologist evaluation of the generated images to assess whether they satisfy clinical desiderata, and (3) training a reward model that predicts human feedback on new samples, which we use to incorporate expert knowledge into the finetuning objective of the diffusion model. Our results show that human feedback significantly improves the quality of synthetic images in terms of fidelity, diversity, utility in downstream applications, and plausibility as evaluated by experts. We also demonstrate that human feedback can teach the model new clinical concepts not annotated in the original training data. Our results demonstrate the value of incorporating human feedback in clinical applications where generative models may struggle to capture extensive domain knowledge from raw data alone.
Shenghuan Sun, Gregory M. Goldgof, Atul J. Butte, Ahmed Alaa 0001
NeurIPS3
2023 Bottom-up and top-down paradigms of artificial intelligence research approaches to healthcare data science using growing real-world big data
abstract
OBJECTIVES: As the real-world electronic health record (EHR) data continue to grow exponentially, novel methodologies involving artificial intelligence (AI) are becoming increasingly applied to enable efficient data-driven learning and, ultimately, to advance healthcare. Our objective is to provide readers with an understanding of evolving computational methods and help in deciding on methods to pursue. TARGET AUDIENCE: The sheer diversity of existing methods presents a challenge for health scientists who are beginning to apply computational methods to their research. Therefore, this tutorial is aimed at scientists working with EHR data who are early entrants into the field of applying AI methodologies. SCOPE: This manuscript describes the diverse and growing AI research approaches in healthcare data science and categorizes them into 2 distinct paradigms, the bottom-up and top-down paradigms to provide health scientists venturing into artificial intelligent research with an understanding of the evolving computational methods and help in deciding on methods to pursue through the lens of real-world healthcare data.
Michelle Wang, Madhumita Sushil, Brenda Y. Miao, Atul J. Butte
J. Am. Medical Informatics Assoc.4
2022 Understanding the Impact of Demographics and Socioeconomic Parameters on Treatment Selection and Utilization in Type 2 Diabetes
Jaysón M. Davidson, Rohit Davidson, Ayan Patel, Atul J. Butte
AMIA4
2022 Deidentifying a Corpus of 100 Million Clinical Text Documents for Information Extraction: Lessons Learned
Lakshmi Radhakrishnan, Gundolf Schenk, Kathlene Muenzen, Boris Oskotsky, Sharat Israni, Atul J. Butte
AMIA6
2022 Training a Transferrable Clinical Language Model from 75 million Notes
Madhumita Sushil, Dana Ludwig, Atul J. Butte, Vivek A. Rudrapatna
AMIA3
2022 Embedding electronic health records onto a knowledge network recognizes prodromal features of multiple sclerosis and predicts diagnosis
abstract
OBJECTIVE: Early identification of chronic diseases is a pillar of precision medicine as it can lead to improved outcomes, reduction of disease burden, and lower healthcare costs. Predictions of a patient's health trajectory have been improved through the application of machine learning approaches to electronic health records (EHRs). However, these methods have traditionally relied on "black box" algorithms that can process large amounts of data but are unable to incorporate domain knowledge, thus limiting their predictive and explanatory power. Here, we present a method for incorporating domain knowledge into clinical classifications by embedding individual patient data into a biomedical knowledge graph. MATERIALS AND METHODS: A modified version of the Page rank algorithm was implemented to embed millions of deidentified EHRs into a biomedical knowledge graph (SPOKE). This resulted in high-dimensional, knowledge-guided patient health signatures (ie, SPOKEsigs) that were subsequently used as features in a random forest environment to classify patients at risk of developing a chronic disease. RESULTS: Our model predicted disease status of 5752 subjects 3 years before being diagnosed with multiple sclerosis (MS) (AUC = 0.83). SPOKEsigs outperformed predictions using EHRs alone, and the biological drivers of the classifiers provided insight into the underpinnings of prodromal MS. CONCLUSION: Using data from EHR as input, SPOKEsigs describe patients at both the clinical and biological levels. We provide a clinical use case for detecting MS up to 5 years prior to their documented diagnosis in the clinic and illustrate the biological features that distinguish the prodromal MS state.
Charlotte A. Nelson, Riley Bove, Atul J. Butte, Sergio Baranzini
J. Am. Medical Informatics Assoc.3
2021 A Generalizable Framework for Cost-Effectiveness Analysis of Antihypertensive Drugs Leveraging Real-World Evidence
Douglas Arneson, Rohit Vashisht, Vivek A. Rudrapatna, Atul J. Butte
AMIA4
2021 Use of Healthcare Information Technology and Data Platforms to Inform Pandemic Response: Perspectives at the Intersection of Academic Healthcare Provider Organizations and Industry Partners
Atul J. Butte, Christopher A. Longhurst, Amy Abernethy, Gretchen Purcell Jackson, Philip R. O. Payne
AMIA1
2021 Systematic identification of ACE2 expression modulators reveals cardiomyopathy as a risk factor for mortality in COVID-19 patients
Navchetan Kaur, Boris Oskotsky, Atul J. Butte, Zicheng Hu
AMIA3
2021 Use of electronic health records to support a public health response to the COVID-19 pandemic in the United States: a perspective from 15 academic medical centers
abstract
Our goal is to summarize the collective experience of 15 organizations in dealing with uncoordinated efforts that result in unnecessary delays in understanding, predicting, preparing for, containing, and mitigating the COVID-19 pandemic in the US. Response efforts involve the collection and analysis of data corresponding to healthcare organizations, public health departments, socioeconomic indicators, as well as additional signals collected directly from individuals and communities. We focused on electronic health record (EHR) data, since EHRs can be leveraged and scaled to improve clinical care, research, and to inform public health decision-making. We outline the current challenges in the data ecosystem and the technology infrastructure that are relevant to COVID-19, as witnessed in our 15 institutions. The infrastructure includes registries and clinical data networks to support population-level analyses. We propose a specific set of strategic next steps to increase interoperability, overall organization, and efficiencies.
Subha Madhavan, Lisa Bastarache, Jeffrey S. Brown, Atul J. Butte, David A. Dorr, Peter J. Embí, Charles P. Friedman, Kevin B. Johnson, Jason H. Moore, Isaac S. Kohane, Philip R. O. Payne, Jessica D. Tenenbaum, Mark G. Weiner, Adam B. Wilcox, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.4
2019 PatientExploreR: an extensible application for dynamic visualization of patient clinical history from electronic health records in the OMOP common data model
abstract
MOTIVATION: Electronic health records (EHRs) are quickly becoming omnipresent in healthcare, but interoperability issues and technical demands limit their use for biomedical and clinical research. Interactive and flexible software that interfaces directly with EHR data structured around a common data model (CDM) could accelerate more EHR-based research by making the data more accessible to researchers who lack computational expertise and/or domain knowledge. RESULTS: We present PatientExploreR, an extensible application built on the R/Shiny framework that interfaces with a relational database of EHR data in the Observational Medical Outcomes Partnership CDM format. PatientExploreR produces patient-level interactive and dynamic reports and facilitates visualization of clinical data without any programming required. It allows researchers to easily construct and export patient cohorts from the EHR for analysis with other software. This application could enable easier exploration of patient-level data for physicians and researchers. PatientExploreR can incorporate EHR data from any institution that employs the CDM for users with approved access. The software code is free and open source under the MIT license, enabling institutions to install and users to expand and modify the application for their own purposes. AVAILABILITY AND IMPLEMENTATION: PatientExploreR can be freely obtained from GitHub: https://github.com/BenGlicksberg/PatientExploreR. We provide instructions for how researchers with approved access to their institutional EHR can use this package. We also release an open sandbox server of synthesized patient data for users without EHR access to explore: http://patientexplorer.ucsf.edu. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Benjamin S. Glicksberg, Boris Oskotsky, Phyllis Thangaraj, Nicholas Giangreco, Marcus A. Badgeley, Kipp W. Johnson, Debajyoti Datta, Vivek A. Rudrapatna, Nadav Rappoport, Mark M. Shervey, Riccardo Miotto, Theodore C. Goldstein, Eugenia Rutenberg, Remi Frazier, Sharat Israni, Rick Larsen, Bethany Percha, Li Li 0062, Joel Dudley, Nicholas P. Tatonetti, Atul J. Butte
Bioinform.22
2019 Robust prediction of clinical outcomes using cytometry data
abstract
MOTIVATION: Flow cytometry and mass cytometry are widely used to diagnose diseases and to predict clinical outcomes. When associating clinical features with cytometry data, traditional analysis methods require cell gating as an intermediate step, leading to information loss and susceptibility to batch effects. Here, we wish to explore an alternative approach that predicts clinical features from cytometry data without the cell-gating step. We also wish to test if such a gating-free approach increases the accuracy and robustness of the prediction. RESULTS: We propose a novel strategy (CytoDx) to predict clinical outcomes using cytometry data without cell gating. Applying CytoDx on real-world datasets allow us to predict multiple types of clinical features. In particular, CytoDx is able to predict the response to influenza vaccine using highly heterogeneous datasets, demonstrating that it is not only accurate but also robust to batch effects and cytometry platforms. AVAILABILITY AND IMPLEMENTATION: CytoDx is available as an R package on Bioconductor (bioconductor.org/packages/CytoDx). Data and scripts for reproducing the results are available on bitbucket.org/zichenghu_ucsf/cytodx_study_code/downloads. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Zicheng Hu, Benjamin S. Glicksberg, Atul J. Butte
Bioinform.3
2016 A survey of current trends in computational drug repositioning
abstract
Computational drug repositioning or repurposing is a promising and efficient tool for discovering new uses from existing drugs and holds the great potential for precision medicine in the age of big data. The explosive growth of large-scale genomic and phenotypic data, as well as data of small molecular compounds with granted regulatory approval, is enabling new developments for computational repositioning. To achieve the shortest path toward new drug indications, advanced data processing and analysis strategies are critical for making sense of these heterogeneous molecular measurements. In this review, we show recent advancements in the critical areas of computational drug repositioning from multiple aspects. First, we summarize available data sources and the corresponding computational repositioning strategies. Second, we characterize the commonly used computational techniques. Third, we discuss validation strategies for repositioning studies, including both computational and experimental methods. Finally, we highlight potential opportunities and use-cases, including a few target areas such as cancers. We conclude with a brief discussion of the remaining challenges in computational drug repositioning.
Jiao Li 0001, Si Zheng 0001, Atul J. Butte, Sanjay Joshua Swamidass, Zhiyong Lu
Briefings Bioinform.4
2016 Immune modulators in disease: integrating knowledge from the biomedical literature and gene expression
abstract
OBJECTIVE: Cytokines play a central role in both health and disease, modulating immune responses and acting as diagnostic markers and therapeutic targets. This work takes a systems-level approach for integration and examination of immune patterns, such as cytokine gene expression with information from biomedical literature, and applies it in the context of disease, with the objective of identifying potentially useful relationships and areas for future research. RESULTS: We present herein the integration and analysis of immune-related knowledge, namely, information derived from biomedical literature and gene expression arrays. Cytokine-disease associations were captured from over 2.4 million PubMed records, in the form of Medical Subject Headings descriptor co-occurrences, as well as from gene expression arrays. Clustering of cytokine-disease co-occurrences from biomedical literature is shown to reflect current medical knowledge as well as potentially novel relationships between diseases. A correlation analysis of cytokine gene expression in a variety of diseases revealed compelling relationships. Finally, a novel analysis comparing cytokine gene expression in different diseases to parallel associations captured from the biomedical literature was used to examine which associations are interesting for further investigation. DISCUSSION: We demonstrate the usefulness of capturing Medical Subject Headings descriptor co-occurrences from biomedical publications in the generation of valid and potentially useful hypotheses. Furthermore, integrating and comparing descriptor co-occurrences with gene expression data was shown to be useful in detecting new, potentially fruitful, and unaddressed areas of research. CONCLUSION: Using integrated large-scale data captured from the scientific literature and experimental data, a better understanding of the immune mechanisms underlying disease can be achieved and applied to research.
Nophar Geifman, Sanchita Bhattacharya, Atul J. Butte
J. Am. Medical Informatics Assoc.3
2016 Constraints on Biological Mechanism from Disease Comorbidity Using Electronic Medical Records and Database of Genetic Variants
abstract
Patterns of disease co-occurrence that deviate from statistical independence may represent important constraints on biological mechanism, which sometimes can be explained by shared genetics. In this work we study the relationship between disease co-occurrence and commonly shared genetic architecture of disease. Records of pairs of diseases were combined from two different electronic medical systems (Columbia, Stanford), and compared to a large database of published disease-associated genetic variants (VARIMED); data on 35 disorders were available across all three sources, which include medical records for over 1.2 million patients and variants from over 17,000 publications. Based on the sources in which they appeared, disease pairs were categorized as having predominant clinical, genetic, or both kinds of manifestations. Confounding effects of age on disease incidence were controlled for by only comparing diseases when they fall in the same cluster of similarly shaped incidence patterns. We find that disease pairs that are overrepresented in both electronic medical record systems and in VARIMED come from two main disease classes, autoimmune and neuropsychiatric. We furthermore identify specific genes that are shared within these disease groups.
Steven C. Bagley, Marina Sirota, Richard O. Chen, Atul J. Butte, Russ B. Altman
PLoS Comput. Biol.4
2015 Recent Advances in Computational Drug Repositioning
Atul J. Butte, Nigam H. Shah, Nicholas P. Tatonetti, Hua Xu 0001
AMIA1
2015 Human Computation of Big Data in Biomedicine: Making STAR annotations for large scale functional characterization of disease
Dexter Hadley, James Pan, Osama M. El-Sayed, Jihad Al-Jabban, Imad Al-Jabban, Tej Azad, Shuaib Raza, Mohamed Hadied, Hyojung Paik, Sanchita Bhattacharya, Marina Sirota, Atul J. Butte
AMIA13
2015 ImmPort: Shared research data for bioinformatics and immunology
abstract
Researchers are applying a wide variety of assay methods to explore the complex and dynamic human immunology system. There are a number of initiatives to encourage sharing of results, the standardization of terminology, and the reanalysis of data to improve reproducibility and ensure scientific rigor. The ImmPort project is an example of a NIH resource to collect, curate, and share immunological research. We describe the ImmPort project resources available to encourage an exploration of the tools and skill sets needed for analyzing shared immunological data, including clinical and mechanistic studies, analysis tools, and example analysis code. There is an ongoing effort to annotate the shared data with standardized terms from ontologies, make these curated data sets available in formats amenable to bioinformatic tools, and provide examples of how to access and analyze the data sets.
Patrick J. Dunn, Elizabeth Thomson, John Campbell 0004, Vincent Desborough, Jeffrey A. Wiser, Henry Schaefer, Sanchita Bhattacharya, Atul J. Butte, Sandra Andorf, Mazen Nasrallah
BIBM9
2014 Investigating the Genetic Architecture of Pulmonary Arterial Hypertension Shared with Other Diseases
Luke Yancy Jr., Atul J. Butte
AMIA2
2013 DTMBIO 2013: international workshop on data and text mining in biomedical informatics
abstract
The organizers of ACM Seventh International Workshop on Data and Text Mining in Biomedical Informatics (DTMBIO 13) are pleased to announce that the seventh DTMBIO will be held in conjunction with CIKM, one of the largest data management conferences. The major interests of DTMBIO are on the state-of-the-art applications of data and text mining on biomedical research problems. DTMBIO 13 will be a forum of discussing and exchanging informatics related techniques and problems in the context of biomedical research.
Atul J. Butte, Doheon Lee, Hua Xu 0001, Min Song 0001
CIKM1
2013 Relating genes to function: identifying enriched transcription factors using the ENCODE ChIP-Seq significance tool
abstract
MOTIVATION: Biological analysis has shifted from identifying genes and transcripts to mapping these genes and transcripts to biological functions. The ENCODE Project has generated hundreds of ChIP-Seq experiments spanning multiple transcription factors and cell lines for public use, but tools for a biomedical scientist to analyze these data are either non-existent or tailored to narrow biological questions. We present the ENCODE ChIP-Seq Significance Tool, a flexible web application leveraging public ENCODE data to identify enriched transcription factors in a gene or transcript list for comparative analyses. IMPLEMENTATION: The ENCODE ChIP-Seq Significance Tool is written in JavaScript on the client side and has been tested on Google Chrome, Apple Safari and Mozilla Firefox browsers. Server-side scripts are written in PHP and leverage R and a MySQL database. The tool is available at http://encodeqt.stanford.edu. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary material is available at Bioinformatics online.
Raymond K. Auerbach, Atul J. Butte
Bioinform.3
2013 Making it personal: translational bioinformatics
abstract
One of the most exciting research areas in Translational Bioinformatics1,2 is related to the redefinition of fundamental notions of what constitutes a ‘disease.’ Nosology, the systematic classification of diseases, dates back to Carl Linnaeus, with the Genera Morborum3 Today, the improvement in our abilities to make molecular measurements related to health and disease has largely driven the revolution towards personalized medicine. For example, in diseases like non-small cell lung cancer or breast cancer, standard-of-care is now including sequencing of genes such as EGFR or quantitating panels of RNA such as those included in Oncotype DX, respectively, to drive therapeutic decisions for new subtypes of patients. While experts, including those at the National Research Council, are seeing the potential of scaling beyond these early case examples towards redefining our entire nosology,4 it is in the field of cancer where personalized or precision medicine has had best traction. It is no coincidence that many contributions to this special issue of JAMIA focus on cancer. Personalized medicine, also known as precision medicine, has often been equated with the use of molecular measurements to characterize disease. The special feature in this issue of JAMIA challenges this limited view.
Atul J. Butte, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.1
2012 Data-driven integration of epidemiological and toxicological data to select candidate interacting genes and environmental factors in association with disease
abstract
MOTIVATION: Complex diseases, such as Type 2 Diabetes Mellitus (T2D), result from the interplay of both environmental and genetic factors. However, most studies investigate either the genetics or the environment and there are a few that study their possible interaction in context of disease. One key challenge in documenting interactions between genes and environment includes choosing which of each to test jointly. Here, we attempt to address this challenge through a data-driven integration of epidemiological and toxicological studies. Specifically, we derive lists of candidate interacting genetic and environmental factors by integrating findings from genome-wide and environment-wide association studies. Next, we search for evidence of toxicological relationships between these genetic and environmental factors that may have an etiological role in the disease. We illustrate our method by selecting candidate interacting factors for T2D.
Chirag J. Patel, Rong Chen 0006, Atul J. Butte
Bioinform.3
2012 Clinical utility of sequence-based genotype compared with that derivable from genotyping arrays
abstract
OBJECTIVE: We investigated the common-disease relevant information obtained from sequencing compared with that reported from genotyping arrays. MATERIALS AND METHODS: Using 187 publicly available individual human genomes, we constructed genomic disease risk summaries based on 55 common diseases with reported gene-disease associations in the research literature using two different risk models, one based on the product of likelihood ratios and the other on the allelic variant with the maximum associated disease risk. We also constructed risk profiles based on the single nucleotide polymorphisms (SNPs) of these individuals that could be measured or imputed from two common genotyping array platforms. RESULTS: We show that the model risk predictions derived from sequencing differ substantially from those obtained from the SNPs measured on commercially available genotyping arrays for several different non-monogenic diseases, although high density genotyping arrays give identical results for many diseases. CONCLUSIONS: Our approach may be used to compare the ability of different platforms to probe known genetic risks disease by disease.
Alexander A. Morgan, Rong Chen 0006, Atul J. Butte
J. Am. Medical Informatics Assoc.3
2012 Multiplex meta-analysis of RNA expression to identify genes with variants associated with immune dysfunction
abstract
OBJECTIVE: We demonstrate a genome-wide method for the integration of many studies of gene expression of phenotypically similar disease processes, a method of multiplex meta-analysis. We use immune dysfunction as an example disease process. DESIGN: We use a heterogeneous collection of datasets across human and mice samples from a range of tissues and different forms of immunodeficiency. We developed a method integrating Tibshirani's modified t-test (SAM) is used to interrogate differential expression within a study and Fisher's method for omnibus meta-analysis to identify differentially expressed genes across studies. The ability of this overall gene expression profile to prioritize disease associated genes is evaluated by comparing against the results of a recent genome wide association study for common variable immunodeficiency (CVID). RESULTS: Our approach is able to prioritize genes associated with immunodeficiency in general (area under the ROC curve = 0.713) and CVID in particular (area under the ROC curve = 0.643). CONCLUSIONS: This approach may be used to investigate a larger range of failures of the immune system. Our method may be extended to other disease processes, using RNA levels to prioritize genes likely to contain disease associated DNA variants.
Alexander A. Morgan, Vasilios J. Pyrgos, Kari C. Nadeau, Peter R. Williamson, Atul J. Butte
J. Am. Medical Informatics Assoc.5
2012 Ten Years of Pathway Analysis: Current Approaches and Outstanding Challenges
abstract
Pathway analysis has become the first choice for gaining insight into the underlying biology of differentially expressed genes and proteins, as it reduces complexity and has increased explanatory power. We discuss the evolution of knowledge base-driven pathway analysis over its first decade, distinctly divided into three generations. We also discuss the limitations that are specific to each generation, and how they are addressed by successive generations of methods. We identify a number of annotation challenges that must be addressed to enable development of the next generation of pathway analysis methods. Furthermore, we identify a number of methodological challenges that the next generation of methods must tackle to take advantage of the technological advances in genomics and proteomics in order to improve specificity, sensitivity, and relevance of pathway analysis.
Purvesh Khatri, Marina Sirota, Atul J. Butte
PLoS Comput. Biol.3
2012 Integrative Approach to Pain Genetics Identifies Pain Sensitivity Loci across Diseases
abstract
Identifying human genes relevant for the processing of pain requires difficult-to-conduct and expensive large-scale clinical trials. Here, we examine a novel integrative paradigm for data-driven discovery of pain gene candidates, taking advantage of the vast amount of existing disease-related clinical literature and gene expression microarray data stored in large international repositories. First, thousands of diseases were ranked according to a disease-specific pain index (DSPI), derived from Medical Subject Heading (MESH) annotations in MEDLINE. Second, gene expression profiles of 121 of these human diseases were obtained from public sources. Third, genes with expression variation significantly correlated with DSPI across diseases were selected as candidate pain genes. Finally, selected candidate pain genes were genotyped in an independent human cohort and prospectively evaluated for significant association between variants and measures of pain sensitivity. The strongest signal was with rs4512126 (5q32, ABLIM3, P = 1.3×10⁻¹⁰) for the sensitivity to cold pressor pain in males, but not in females. Significant associations were also observed with rs12548828, rs7826700 and rs1075791 on 8q22.2 within NCALD (P = 1.7×10⁻⁴, 1.8×10⁻⁴, and 2.2×10⁻⁴ respectively). Our results demonstrate the utility of a novel paradigm that integrates publicly available disease-specific gene expression data with clinical data curated from MEDLINE to facilitate the discovery of pain-relevant genes. This data-derived list of pain gene candidates enables additional focused and efficient biological studies validating additional candidates.
David Ruau, Joel Dudley, Rong Chen 0006, Nicholas G. Phillips, Gary E. Swan, Laura Lazzeroni, J. David Clark, Atul J. Butte, Martin S. Angst
PLoS Comput. Biol.8
2011 Exploiting drug-disease relationships for computational drug repositioning
abstract
Finding new uses for existing drugs, or drug repositioning, has been used as a strategy for decades to get drugs to more patients. As the ability to measure molecules in high-throughput ways has improved over the past decade, it is logical that such data might be useful for enabling drug repositioning through computational methods. Many computational predictions for new indications have been borne out in cellular model systems, though extensive animal model and clinical trial-based validation are still pending. In this review, we show that computational methods for drug repositioning can be classified in two axes: drug based, where discovery initiates from the chemical perspective, or disease based, where discovery initiates from the clinical perspective of disease or its pathology. Newer algorithms for computational drug repositioning will likely span these two axes, will take advantage of newer types of molecular measurements, and will certainly play a role in reducing the global burden of disease.
Joel Dudley, Tarangini Deshpande, Atul J. Butte
Briefings Bioinform.3
2011 ProfileChaser: searching microarray repositories based on genome-wide patterns of differential expression
abstract
SUMMARY: We introduce ProfileChaser, a web server that allows for querying the Gene Expression Omnibus based on genome-wide patterns of differential expression. Using a novel, content-based approach, ProfileChaser retrieves expression profiles that match the differentially regulated transcriptional programs in a user-supplied experiment. This analysis identifies statistical links to similar expression experiments from the vast array of publicly available data on diseases, drugs, phenotypes and other experimental conditions. AVAILABILITY: http://profilechaser.stanford.edu CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jesse M. Engreitz, Rong Chen 0006, Alexander A. Morgan, Joel Dudley, Rohan Mallelwar, Atul J. Butte
Bioinform.6
2011 Computationally translating molecular discoveries into tools for medicine: translational bioinformatics articles now featured in JAMIA
abstract
This year marks the 15th anniversary of the invention of the gene expression microarray. As mRNA transcripts serve as the blueprint within cells for making proteins, measuring mRNA levels was seen as an accurate and manageable way to investigate cell and tissue processes. Those earliest microarrays in 1995 could measure 48 transcripts in parallel in plants,1 but within 1 year were scaled up to measure more than 1000 transcripts including those in human tissues. Today, these microarrays are essentially commodity items, commonly used to study human health and disease in hospitals and academic institutions, as well as in the biotechnology and pharmaceutical industry. While tens of thousands of publications have already been published referencing microarrays, this is just the start. Similar arrays are already used to probe genetic differences in DNA, but even these will soon be supplanted by whole genome sequencing, where we can expect all three billion human base pairs to be sequenced for a few thousand dollars. The exponential decrease in costs for whole genome sequencing has been described as going beyond the decline we are used to from Moore's law.2 These enormous amounts of molecular data make it clear that there is a pressing need for computational methods to analyze and interpret them. Molecular data have never been foreign to the pages of JAMIA or AMIA Symposia. The second volume of JAMIA back in 1995 contained an article describing how the internet and newly introduced world wide web could be used to facilitate genome sequencing efforts across two academic genome centers and introduced concepts like yeast artificial chromosomes, sequence tagged sites and contigs.3 The 2002 AMIA Fall Symposium suggested bioinformatics and medical informatics could and should be part of the same discipline of biomedical informatics. While not every publication in bioinformatics is likely relevant to AMIA members, it is important to track those applying bioinformatics to human health and disease. Translational bioinformatics can be defined as ‘the development of storage, analytic, and interpretive methods to optimize the transformation of increasingly voluminous biomedical data into proactive, predictive, preventive, and participatory health.’4 Indeed, translational bioinformatics has been a core strategic area of importance for AMIA since 2008. Today, as more hospital and academic medical centers embrace molecular measurements for diagnosis and planning therapies, we recognize that the community of JAMIA readers will need to keep abreast of new developments and applications of translational bioinformatics.5 In the past few years, however, it has been increasingly hard to find articles on translational bioinformatics in JAMIA. However, with this month's issue, we hope to start reversing the trend. Starting with a Perspective on from Neil Sarkar and colleagues (see page 354), we are highlighting five manuscripts in the field of translational bioinformatics, and ‘opening the doors' for submissions from investigators and authors in this field.6 Recognizing the continued growth in bioinformatics, especially as related to human health and disease, in 2009 AMIA initiated the annual Summit on Translational Bioinformatics as a new annual meeting to address the growing need for a scientific conference to present and discuss developments and the application of methods in this field. At the 2011 Summit, several authors of top-ranked submitted poster abstracts were invited to expand their submissions into full manuscripts, which were then evaluated by JAMIA reviewers. One such manuscript appears in this month's issue. Hua Xu and colleagues (see page 387) show how dosing details locked within free-text electronic health records could be found using natural language processing, thus enabling a gene–dosing association study for warfarin dosing at Vanderbilt University.7 This work serves as a premier model of the type of translational research possible when bioinformatics and medical informatics investigators truly collaborate. Few institutions have the considerable resources of Vanderbilt University, which has a DNA biobank linked to a de-identified electronic health record subset, but there are a few, notably the Mayo Clinic. In a manuscript this month, Pathak and colleagues (see page 376) show how standardized representations of clinical characteristics extracted from institutional electronic health records can be integrated to enable large scale genetics studies between Vanderbilt University and the Mayo Clinic.8 As the genetic risk of disease is now theorized to result from a combination of rare DNA variants,9 larger cross-country cohorts like these will be needed to identify these variants. Integration of information on patients, samples, or data is indeed a common theme in translational bioinformatics. David Foran and colleagues (see page 403) highlight a new federated software system that can handle another type of biobank for pathology samples.10 Processed into histology cores and distributed on a tissue microarray, these samples are proving to be invaluable for the high-throughput query of specific proteins. James Chen and colleagues (see page 392) show how comparing the genomic differences found across publicly available data from previous prostate cancer studies can yield a single core ‘signature’ that can distinguish prostate cancer patients with better prognosis from those with worse prognosis.11 Translational bioinformatics grew out of the work of a small but cohesive group of researchers who bridged the gap between computational biology and medicine. It is with sad remembrance that we note the passing of Marco Ramoni, a pioneering spirit who was one of the first Track Chairs for the AMIA Summit on Translational Bioinformatics. So that his contributions to AMIA and the field of translational bioinformatics will not be forgotten, this year the Board of Directors established the Marco Ramoni Best Paper Award, to be presented each year at the AMIA Summit on Translational Bioinformatics. The manuscript from the award winner for 2011, Wei Wei, appears in this issue of JAMIA (see page 370).12 Selected by an external review panel chosen by the Chair of the Scientific Program Committee and again peer reviewed by JAMIA reviewers, it is slightly ironic that the award winning manuscript is on the application of Bayes' theorem, coincidentally similar to Dr Ramoni's own work, which is summarized by Kohane and Szolovits (see page 367) in an invited academic tribute to our dear colleague.13 Next year will see the fifth Summit on Translational Bioinformatics. With the approaching changes in the amount and diversity of datasets discussed above, we anticipate that data-centric approaches that compute on massive amounts of data (often called ‘big data’14) to identify patterns and make clinically relevant predictions will be increasingly common in translational bioinformatics. In anticipation, the 2012 Summit on Translational Bioinformatics will have four tracks focusing on research that take us from base pairs to the bedside,15 with a particular emphasis on the clinical implications of mining massive datasets, and bridging the latest multimodal measurement technologies using the large amounts of electronic healthcare data that are increasingly available. We invite readers to submit extended (10-page) submissions for the Summit before the deadline of August 15. The top papers will be published in JAMIA after peer review and the editorial office will make all efforts to have them available online first by the time the Summit takes place on March 19–21, 2012 in San Francisco. In closing, we have continued to note arguments over how much translational bioinformatics informatics investigators and professionals really need to know. To answer this, we must consider that we are entering a decade where tens of thousands of people have already obtained samplings of their own DNA sequences from consumer genomics companies,16 with over 700 000 RNA microarray measurements already available to the public,17,18 and we expect that 30 000 people will have their whole genome sequenced this year alone.19 We must not keep assuming that if we just build and provide the right generalized tools and methods, others will take them and use them the right way to improve healthcare—the field is over-saturated with tools already. If we best understand the tools and methods we build, we need to be the first to actually use those tools, and show the world what can be achieved. Instead of discussing how relevant translational bioinformatics is, we need to argue that biomedical informatics is the only field in biomedicine that is ready to revolutionize human health and healthcare using these tools and measurements. AJB is funded by the US National Library of Medicine (R01 LM009719) and the Lucile Packard Foundation for Children's Health. NHS is funded by the US National Institute of Health Roadmap (U54 HG004028). None. Commissioned; internally peer reviewed.
Atul J. Butte, Nigam H. Shah
J. Am. Medical Informatics Assoc.1
2011 Translational bioinformatics: linking knowledge across biological and clinical realms
abstract
Nearly a decade since the completion of the first draft of the human genome, the biomedical community is positioned to usher in a new era of scientific inquiry that links fundamental biological insights with clinical knowledge. Accordingly, holistic approaches are needed to develop and assess hypotheses that incorporate genotypic, phenotypic, and environmental knowledge. This perspective presents translational bioinformatics as a discipline that builds on the successes of bioinformatics and health informatics for the study of complex diseases. The early successes of translational bioinformatics are indicative of the potential to achieve the promise of the Human Genome Project for gaining deeper insights to the genetic underpinnings of disease and progress toward the development of a new generation of therapies.
Indra Neil Sarkar, Atul J. Butte, Yves A. Lussier, Peter Tarczy-Hornoch, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.2
2011 Comparison of automated and human assignment of MeSH terms on publicly-available molecular datasets
abstract
Publicly available molecular datasets can be used for independent verification or investigative repurposing, but depends on the presence, consistency and quality of descriptive annotations. Annotation and indexing of molecular datasets using well-defined controlled vocabularies or ontologies enables accurate and systematic data discovery, yet the majority of molecular datasets available through public data repositories lack such annotations. A number of automated annotation methods have been developed; however few systematic evaluations of the quality of annotations supplied by application of these methods have been performed using annotations from standing public data repositories. Here, we compared manually-assigned Medical Subject Heading (MeSH) annotations associated with experiments by data submitters in the PRoteomics IDEntification (PRIDE) proteomics data repository to automated MeSH annotations derived through the National Center for Biomedical Ontology Annotator and National Library of Medicine MetaMap programs. These programs were applied to free-text annotations for experiments in PRIDE. As many submitted datasets were referenced in publications, we used the manually curated MeSH annotations of those linked publications in MEDLINE as "gold standard". Annotator and MetaMap exhibited recall performance 3-fold greater than that of the manual annotations. We connected PRIDE experiments in a network topology according to shared MeSH annotations and found 373 distinct clusters, many of which were found to be biologically coherent by network analysis. The results of this study suggest that both Annotator and MetaMap are capable of annotating public molecular datasets with a quality comparable, and often exceeding, that of the actual data submitters, highlighting a continuous need to improve and apply automated methods to molecular datasets in public data repositories to maximize their value and utility.
David Ruau, Michael Mbagwu, Joel Dudley, Vijay Krishnan, Atul J. Butte
J. Biomed. Informatics5
2010 Latent physiological factors of complex human diseases revealed by independent component analysis of clinarrays
abstract
BACKGROUND: Diagnosis and treatment of patients in the clinical setting is often driven by known symptomatic factors that distinguish one particular condition from another. Treatment based on noticeable symptoms, however, is limited to the types of clinical biomarkers collected, and is prone to overlooking dysfunctions in physiological factors not easily evident to medical practitioners. We used a vector-based representation of patient clinical biomarkers, or clinarrays, to search for latent physiological factors that underlie human diseases directly from clinical laboratory data. Knowledge of these factors could be used to improve assessment of disease severity and help to refine strategies for diagnosis and monitoring disease progression. RESULTS: Applying Independent Component Analysis on clinarrays built from patient laboratory measurements revealed both known and novel concomitant physiological factors for asthma, types 1 and 2 diabetes, cystic fibrosis, and Duchenne muscular dystrophy. Serum sodium was found to be the most significant factor for both type 1 and type 2 diabetes, and was also significant in asthma. TSH3, a measure of thyroid function, and blood urea nitrogen, indicative of kidney function, were factors unique to type 1 diabetes respective to type 2 diabetes. Platelet count was significant across all the diseases analyzed. CONCLUSIONS: The results demonstrate that large-scale analyses of clinical biomarkers using unsupervised methods can offer novel insights into the pathophysiological basis of human disease, and suggest novel clinical utility of established laboratory measurements.
David P. Chen, Joel Dudley, Atul J. Butte
BMC Bioinform.3
2010 Content-based microarray search using differential expression profiles
abstract
BACKGROUND: With the expansion of public repositories such as the Gene Expression Omnibus (GEO), we are rapidly cataloging cellular transcriptional responses to diverse experimental conditions. Methods that query these repositories based on gene expression content, rather than textual annotations, may enable more effective experiment retrieval as well as the discovery of novel associations between drugs, diseases, and other perturbations. RESULTS: We develop methods to retrieve gene expression experiments that differentially express the same transcriptional programs as a query experiment. Avoiding thresholds, we generate differential expression profiles that include a score for each gene measured in an experiment. We use existing and novel dimension reduction and correlation measures to rank relevant experiments in an entirely data-driven manner, allowing emergent features of the data to drive the results. A combination of matrix decomposition and p-weighted Pearson correlation proves the most suitable for comparing differential expression profiles. We apply this method to index all GEO DataSets, and demonstrate the utility of our approach by identifying pathways and conditions relevant to transcription factors Nanog and FoxO3. CONCLUSIONS: Content-based gene expression search generates relevant hypotheses for biological inquiry. Experiments across platforms, tissue types, and protocols inform the analysis of new datasets.
Jesse M. Engreitz, Alexander A. Morgan, Joel Dudley, Rong Chen 0006, Rahul Thathoo, Russ B. Altman, Atul J. Butte
BMC Bioinform.7
2010 Comparison of multiplex meta analysis techniques for understanding the acute rejection of solid organ transplants
abstract
BACKGROUND: Combining the results of studies using highly parallelized measurements of gene expression such as microarrays and RNAseq offer unique challenges in meta analysis. Motivated by a need for a deeper understanding of organ transplant rejection, we combine the data from five separate studies to compare acute rejection versus stability after solid organ transplantation, and use this data to examine approaches to multiplex meta analysis. RESULTS: We demonstrate that a commonly used parametric effect size estimate approach and a commonly used non-parametric method give very different results in prioritizing genes. The parametric method providing a meta effect estimate was superior at ranking genes based on our gold-standard of identifying immune response genes in the transplant rejection datasets. CONCLUSION: Different methods of multiplex analysis can give substantially different results. The method which is best for any given application will likely depend on the particular domain, and it remains for future work to see if any one method is consistently better at identifying important biological signal across gene expression experiments.
Alexander A. Morgan, Purvesh Khatri, Richard Hayden Jones, Minnie M. Sarwal, Atul J. Butte
BMC Bioinform.5
2010 An integrative method for scoring candidate genes from association studies: application to warfarin dosing
abstract
BACKGROUND: A key challenge in pharmacogenomics is the identification of genes whose variants contribute to drug response phenotypes, which can include severe adverse effects. Pharmacogenomics GWAS attempt to elucidate genotypes predictive of drug response. However, the size of these studies has severely limited their power and potential application. We propose a novel knowledge integration and SNP aggregation approach for identifying genes impacting drug response. Our SNP aggregation method characterizes the degree to which uncommon alleles of a gene are associated with drug response. We first use pre-existing knowledge sources to rank pharmacogenes by their likelihood to affect drug response. We then define a summary score for each gene based on allele frequencies and train linear and logistic regression classifiers to predict drug response phenotypes. RESULTS: We applied our method to a published warfarin GWAS data set comprising 181 individuals. We find that our method can increase the power of the GWAS to identify both VKORC1 and CYP2C9 as warfarin pharmacogenes, where the original analysis had only identified VKORC1. Additionally, we find that our method can be used to discriminate between low-dose (AUROC=0.886) and high-dose (AUROC=0.764) responders. CONCLUSIONS: Our method offers a new route for candidate pharmacogene discovery from pharmacogenomics GWAS, and serves as a foundation for future work in methods for predictive pharmacogenomics.
Nicholas P. Tatonetti, Joel Dudley, Hersh Sagreiya, Atul J. Butte, Russ B. Altman
BMC Bioinform.4
2010 Validating pathophysiological models of aging using clinical electronic medical records
David P. Chen, Alexander A. Morgan, Atul J. Butte
J. Biomed. Informatics3
2010 Current methodologies for translational bioinformatics
Yves A. Lussier, Atul J. Butte, Lawrence Hunter
J. Biomed. Informatics2
2010 Differentially Expressed RNA from Public Microarray Data Identifies Serum Protein Biomarkers for Cross-Organ Transplant Rejection and Other Conditions
abstract
Serum proteins are routinely used to diagnose diseases, but are hard to find due to low sensitivity in screening the serum proteome. Public repositories of microarray data, such as the Gene Expression Omnibus (GEO), contain RNA expression profiles for more than 16,000 biological conditions, covering more than 30% of United States mortality. We hypothesized that genes coding for serum- and urine-detectable proteins, and showing differential expression of RNA in disease-damaged tissues would make ideal diagnostic protein biomarkers for those diseases. We showed that predicted protein biomarkers are significantly enriched for known diagnostic protein biomarkers in 22 diseases, with enrichment significantly higher in diseases for which at least three datasets are available. We then used this strategy to search for new biomarkers indicating acute rejection (AR) across different types of transplanted solid organs. We integrated three biopsy-based microarray studies of AR from pediatric renal, adult renal and adult cardiac transplantation and identified 45 genes upregulated in all three. From this set, we chose 10 proteins for serum ELISA assays in 39 renal transplant patients, and discovered three that were significantly higher in AR. Interestingly, all three proteins were also significantly higher during AR in the 63 cardiac transplant recipients studied. Our best marker, serum PECAM1, identified renal AR with 89% sensitivity and 75% specificity, and also showed increased expression in AR by immunohistochemistry in renal, hepatic and cardiac transplant biopsies. Our results demonstrate that integrating gene expression microarray measurements from disease samples and even publicly-available data sets can be a powerful, fast, and cost-effective strategy for the discovery of new diagnostic serum protein biomarkers.
Rong Chen 0006, Tara K. Sigdel, Li Li 0062, Neeraja Kambham, Joel Dudley, Szu-chuan Hsieh, R. Bryan Klassen, Amery Chen, Tuyen Caohuu, Alexander A. Morgan, Hannah A. Valantine, Kiran K. Khush, Minnie M. Sarwal, Atul J. Butte
PLoS Comput. Biol.14
2010 Network-Based Elucidation of Human Disease Similarities Reveals Common Functional Modules Enriched for Pluripotent Drug Targets
abstract
Current work in elucidating relationships between diseases has largely been based on pre-existing knowledge of disease genes. Consequently, these studies are limited in their discovery of new and unknown disease relationships. We present the first quantitative framework to compare and contrast diseases by an integrated analysis of disease-related mRNA expression data and the human protein interaction network. We identified 4,620 functional modules in the human protein network and provided a quantitative metric to record their responses in 54 diseases leading to 138 significant similarities between diseases. Fourteen of the significant disease correlations also shared common drugs, supporting the hypothesis that similar diseases can be treated by the same drugs, allowing us to make predictions for new uses of existing drugs. Finally, we also identified 59 modules that were dysregulated in at least half of the diseases, representing a common disease-state "signature". These modules were significantly enriched for genes that are known to be drug targets. Interestingly, drugs known to target these genes/proteins are already known to treat significantly more diseases than drugs targeting other genes/proteins, highlighting the importance of these core modules as prime therapeutic opportunities.
Silpa Suthram, Joel Dudley, Annie P. Chiang, Rong Chen 0006, Trevor J. Hastie, Atul J. Butte
PLoS Comput. Biol.6
2009 A Classifier-based approach to identify genetic similarities between diseases
abstract
MOTIVATION: Genome-wide association studies are commonly used to identify possible associations between genetic variations and diseases. These studies mainly focus on identifying individual single nucleotide polymorphisms (SNPs) potentially linked with one disease of interest. In this work, we introduce a novel methodology that identifies similarities between diseases using information from a large number of SNPs. We separate the diseases for which we have individual genotype data into one reference disease and several query diseases. We train a classifier that distinguishes between individuals that have the reference disease and a set of control individuals. This classifier is then used to classify the individuals that have the query diseases. We can then rank query diseases according to the average classification of the individuals in each disease set, and identify which of the query diseases are more similar to the reference disease. We repeat these classification and comparison steps so that each disease is used once as reference disease. RESULTS: We apply this approach using a decision tree classifier to the genotype data of seven common diseases and two shared control sets provided by the Wellcome Trust Case Control Consortium. We show that this approach identifies the known genetic similarity between type 1 diabetes and rheumatoid arthritis, and identifies a new putative similarity between bipolar disease and hypertension.
Marc A. Schaub, Irene M. Kaplow, Marina Sirota, Chuong B. Do, Atul J. Butte, Serafim Batzoglou
Bioinform.5
2009 Selected proceedings of the First Summit on Translational Bioinformatics 2008
Atul J. Butte, Indra Neil Sarkar, Marco Ramoni, Yves A. Lussier, Olga G. Troyanskaya
BMC Bioinform.1
2009 Infection in the intensive care unit alters physiological networks
abstract
BACKGROUND: Physicians use clinical and physiological data to treat patients every day, and it is essential for treating a patient appropriately. However, medical sources of clinical physiological data are only now starting to find use in bioinformatics research. RESULTS: We collected 29 types of physiological and clinical data on a minute-by-minute basis from trauma patients in the intensive care unit along with whether they contracted an infection during their stay. Dividing the patients into two groups based on this criterion, we determined that the correlational network amongst pairs of physiological variables changes based on whether the patient contracted an infection. CONCLUSION: Examining the variable pairs with the largest change in correlation across groups reveals potential changes in the way our treatments affect the patient's physiology and in how our bodies react to physiological insults. These findings highlight the usefulness of physiological informatics and suggest new relationships to study while also validating previously reported relationships.
Adam D. Grossman, Mitchell J. Cohen, Geoffrey T. Manley, Atul J. Butte
BMC Bioinform.4
2009 The "etiome": identification and clustering of human disease etiological factors
abstract
BACKGROUND: Both genetic and environmental factors contribute to human diseases. Most common diseases are influenced by a large number of genetic and environmental factors, most of which individually have only a modest effect on the disease. Though genetic contributions are relatively well characterized for some monogenetic diseases, there has been no effort at curating the extensive list of environmental etiological factors. RESULTS: From a comprehensive search of the MeSH annotation of MEDLINE articles, we identified 3,342 environmental etiological factors associated with 3,159 diseases. We also identified 1,100 genes associated with 1,034 complex diseases from the NIH Genetic Association Database (GAD), a database of genetic association studies. 863 diseases have both genetic and environmental etiological factors available. Integrating genetic and environmental factors results in the "etiome", which we define as the comprehensive compendium of disease etiology. Clustering of environmental factors may alert clinicians of the risks of added exposures, or synergy in interventions to alter these factors. Clustering of both genetic and environmental etiological factors puts genes in the context of environment in a quantitative manner. CONCLUSION: In this paper, we obtained a comprehensive list of associations between disease and environmental factors using MeSH annotation of MEDLINE articles. It serves as a summary of current knowledge between etiological factors and diseases. By combining the environmental etiological factors and genetic factors from GAD, we computed the "etiome" profile for 863 diseases. Comparing diseases across these profiles may have utility for clinical medicine, basic science research, and population-based science.
Yueyi I. Liu, Paul H. Wise, Atul J. Butte
BMC Bioinform.3
2009 Ontology-driven indexing of public datasets for translational bioinformatics
abstract
The volume of publicly available genomic scale data is increasing. Genomic datasets in public repositories are annotated with free-text fields describing the pathological state of the studied sample. These annotations are not mapped to concepts in any ontology, making it difficult to integrate these datasets across repositories. We have previously developed methods to map text-annotations of tissue microarrays to concepts in the NCI thesaurus and SNOMED-CT. In this work we generalize our methods to map text annotations of gene expression datasets to concepts in the UMLS. We demonstrate the utility of our methods by processing annotations of datasets in the Gene Expression Omnibus. We demonstrate that we enable ontology-based querying and integration of tissue and gene expression microarray data. We enable identification of datasets on specific diseases across both repositories. Our approach provides the basis for ontology-driven data integration for translational research on gene and protein expression data. Based on this work we have built a prototype system for ontology based annotation and indexing of biomedical data. The system processes the text metadata of diverse resource elements such as gene expression data sets, descriptions of radiology images, clinical-trial reports, and PubMed article abstracts to annotate and index them with concepts from appropriate ontologies. The key functionality of this system is to enable users to locate biomedical data resources related to particular ontology concepts.
Nigam H. Shah, Clément Jonquet, Annie P. Chiang, Atul J. Butte, Rong Chen 0006, Mark A. Musen
BMC Bioinform.4
2009 Use of Bayesian networks to probabilistically model and improve the likelihood of validation of microarray findings by RT-PCR
Sangeeta B. English, Shou-Ching Shih, Marco Ramoni, Lois E. Smith, Atul J. Butte
J. Biomed. Informatics5
2009 A Quick Guide for Developing Effective Bioinformatics Programming Skills
abstract
Bioinformatics programming skills are becoming a necessity across many facets of biology and medicine, owed in part to the continuing explosion of biological data
Joel Dudley, Atul J. Butte
PLoS Comput. Biol.2
2008 GeneChaser: Identifying all biological and clinical conditions in which genes of interest are differentially expressed
abstract
BACKGROUND: The amount of gene expression data in the public repositories, such as NCBI Gene Expression Omnibus (GEO) has grown exponentially, and provides a gold mine for bioinformaticians, but has not been easily accessible by biologists and clinicians. RESULTS: We developed an automated approach to annotate and analyze all GEO data sets, including 1,515 GEO data sets from 231 microarray types across 42 species, and performed 12,658 group versus group comparisons of 24 GEO-specified types. We then built GeneChaser, a web server that enables biologists and clinicians without bioinformatics skills to easily identify biological and clinical conditions in which a gene or set of genes was differentially expressed. GeneChaser displays these conditions in graphs, gives statistical comparisons, allows sort/filter functions and provides access to the original studies.We performed a single gene search for Nanog and a multiple gene search for Nanog, Oct4, Sox2 and LIN28, confirmed their roles in embryonic stem cell development, identified several drugs that regulate their expression, and suggested their potential roles in sex determination, abnormal sperm morphology, malaria infection, and cancer. CONCLUSION: We demonstrated that GeneChaser is a powerful tool to elucidate information on function, transcriptional regulation, drug-response and clinical implications for genes of interest.
Rong Chen 0006, Rohan Mallelwar, Ajit Thosar, Shivkumar Venkatasubrahmanyam, Atul J. Butte
BMC Bioinform.5
2008 Viewpoint Paper: Translational Bioinformatics: Coming of Age
abstract
The American Medical Informatics Association (AMIA) recently augmented the scope of its activities to encompass translational bioinformatics as a third major domain of informatics. The AMIA has defined translational bioinformatics as "... the development of storage, analytic, and interpretive methods to optimize the transformation of increasingly voluminous biomedical data into proactive, predictive, preventative, and participatory health." In this perspective, I will list eight reasons why this is an excellent time to be studying translational bioinformatics, including the significant increase in funding opportunities available for informatics from the United States National Institutes of Health, and the explosion of publicly-available data sets of molecular measurements. I end with the significant challenges we face in building a community of future investigators in Translational Bioinformatics.
Atul J. Butte
J. Am. Medical Informatics Assoc.1
2007 Clinical Arrays of Laboratory Measures, or "Clinarrays", Built from an Electronic Health Record Enable Disease Subtyping by Severity
David P. Chen, Susan C. Weber, Philip S. Constantinou, Todd A. Ferris, Henry J. Lowe, Atul J. Butte
AMIA6
2007 Methodologies for Extracting Functional Pharmacogenomic Experiments from International Repository
Yi-An Lin, Annie P. Chiang, Ray Lin, Peggy Yao, Rong Chen 0006, Atul J. Butte
AMIA6
2007 Evaluation and integration of 49 genome-wide experiments and the prediction of previously unknown obesity-related genes
abstract
MOTIVATION: Genome-wide experiments only rarely show resounding success in yielding genes associated with complex polygenic disorders. We evaluate 49 obesity-related genome-wide experiments with publicly available findings including microarray, genetics, proteomics and gene knock-down from human, mouse, rat and worm, in terms of their ability to rediscover a comprehensive set of genes previously found to be causally associated or having variants associated with obesity. RESULTS: Individual experiments show poor predictive ability for rediscovering known obesity-associated genes. We show that intersecting the results of experiments significantly improves the sensitivity, specificity and precision of the prediction of obesity-associated genes. We create an integrative model that statistically significantly outperforms all 49 individual genome-wide experiments. We find that genes known to be associated with obesity are significantly implicated in more obesity-related experiments and use this to provide a list of genes that we predict to have the highest likelihood of association for obesity. The approach described here can include any number and type of genome-wide experiments and might be useful for other complex polygenic disorders as well.
Sangeeta B. English, Atul J. Butte
Bioinform.2
2006 Finding Disease-Related Genomic Experiments Within an International Repository: First Steps in Translational Bioinformatics
Atul J. Butte, Rong Chen 0006
AMIA1
2005 Systematic survey reveals general applicability of "guilt-by-association" within gene coexpression networks
abstract
BACKGROUND: Biological processes are carried out by coordinated modules of interacting molecules. As clustering methods demonstrate that genes with similar expression display increased likelihood of being associated with a common functional module, networks of coexpressed genes provide one framework for assigning gene function. This has informed the guilt-by-association (GBA) heuristic, widely invoked in functional genomics. Yet although the idea of GBA is accepted, the breadth of GBA applicability is uncertain. RESULTS: We developed methods to systematically explore the breadth of GBA across a large and varied corpus of expression data to answer the following question: To what extent is the GBA heuristic broadly applicable to the transcriptome and conversely how broadly is GBA captured by a priori knowledge represented in the Gene Ontology (GO)? Our study provides an investigation of the functional organization of five coexpression networks using data from three mammalian organisms. Our method calculates a probabilistic score between each gene and each Gene Ontology category that reflects coexpression enrichment of a GO module. For each GO category we use Receiver Operating Curves to assess whether these probabilistic scores reflect GBA. This methodology applied to five different coexpression networks demonstrates that the signature of guilt-by-association is ubiquitous and reproducible and that the GBA heuristic is broadly applicable across the population of nine hundred Gene Ontology categories. We also demonstrate the existence of highly reproducible patterns of coexpression between some pairs of GO categories. CONCLUSION: We conclude that GBA has universal value and that transcriptional control may be more modular than previously realized. Our analyses also suggest that methodologies combining coexpression measurements across multiple genes in a biologically-defined module can aid in characterizing gene function or in characterizing whether pairs of functions operate together.
Cecily J. Wolfe, Isaac S. Kohane, Atul J. Butte
BMC Bioinform.3
2004 Quantifying the relationship between co-expression, co-regulation and gene function
abstract
BACKGROUND: It is thought that genes with similar patterns of mRNA expression and genes with similar functions are likely to be regulated via the same mechanisms. It has been difficult to quantitatively test these hypotheses on a large scale because there has been no general way of determining whether genes share a common regulatory mechanism. Here we use data from a recent genome wide binding analysis in combination with mRNA expression data and existing functional annotations to quantify the likelihood that genes with varying degrees of similarity in mRNA expression profile or function will be bound by a common transcription factor. RESULTS: Genes with strongly correlated mRNA expression profiles are more likely to have their promoter regions bound by a common transcription factor. This effect is present only at relatively high levels of expression similarity. In order for two genes to have a greater than 50% chance of sharing a common transcription factor binder, the correlation between their expression profiles (across the 611 microarrays used in our study) must be greater than 0.84. Genes with similar functional annotations are also more likely to be bound by a common transcription factor. Combining mRNA expression data with functional annotation results in a better predictive model than using either data source alone. CONCLUSIONS: We demonstrate how mRNA expression data and functional annotations can be used together to estimate the probability that genes share a common regulatory mechanism. Existing microarray data and known functional annotations are sufficient to identify only a relatively small percentage of co-regulated genes.
Dominic J. Allocco, Isaac S. Kohane, Atul J. Butte
BMC Bioinform.3
2003 PGAGENE: integrating quantitative gene-specific results from the NHLBI Programs for Genomic Applications
abstract
Abstract Summary: PGAGENE is a web-based gene-specific genomic data search engine, which allows users to search over 5.9 million pieces of collective genetic and genomic data from the NHLBI supported Programs for Genomic Applications. This data includes microarray measurements, SNPs, and mutations, and data may be found using symbols, parts of gene names or products, Affymetrix probe IDs, GenBank accession numbers, UniGene IDs, dbSNP IDs, and others. The PGAGENE indexing agent periodically maps all publicly available gene-specific PGA data onto LocusLink using dynamically generated cross-referencing tables. Availability: http://pgagene.chip.org Contact: [email protected] * To whom correspondence should be addressed.
Kyungjoon Lee, Isaac S. Kohane, Atul J. Butte
Bioinform.3
2003 Reproducibility of gene expression across generations of Affymetrix microarrays
abstract
BACKGROUND: The development of large-scale gene expression profiling technologies is rapidly changing the norms of biological investigation. But the rapid pace of change itself presents challenges. Commercial microarrays are regularly modified to incorporate new genes and improved target sequences. Although the ability to compare datasets across generations is crucial for any long-term research project, to date no means to allow such comparisons have been developed. In this study the reproducibility of gene expression levels across two generations of Affymetrix GeneChips (HuGeneFL and HG-U95A) was measured. RESULTS: Correlation coefficients were computed for gene expression values across chip generations based on different measures of similarity. Comparing the absolute calls assigned to the individual probe sets across the generations found them to be largely unchanged. CONCLUSION: We show that experimental replicates are highly reproducible, but that reproducibility across generations depends on the degree of similarity of the probe sets and the expression level of the corresponding transcript.
Ashish Nimgaonkar, Despina Sanoudou, Atul J. Butte, Judith N. Haslett, Louis M. Kunkel, Alan H. Beggs, Isaac S. Kohane
BMC Bioinform.3
2002 Analysis of matched mRNA measurements from two different microarray technologies
abstract
MOTIVATION: [corrected] The existence of several technologies for measuring gene expression makes the question of cross-technology agreement of measurements an important issue. Cross-platform utilization of data from different technologies has the potential to reduce the need to duplicate experiments but requires corresponding measurements to be comparable. METHODS: A comparison of mRNA measurements of 2895 sequence-matched genes in 56 cell lines from the standard panel of 60 cancer cell lines from the National Cancer Institute (NCI 60) was carried out by calculating correlation between matched measurements and calculating concordance between cluster from two high-throughput DNA microarray technologies, Stanford type cDNA microarrays and Affymetrix oligonucleotide microarrays. RESULTS: In general, corresponding measurements from the two platforms showed poor correlation. Clusters of genes and cell lines were discordant between the two technologies, suggesting that relative intra-technology relationships were not preserved. GC-content, sequence length, average signal intensity, and an estimator of cross-hybridization were found to be associated with the degree of correlation. This suggests gene-specific, or more correctly probe-specific, factors influencing measurements differently in the two platforms, implying a poor prognosis for a broad utilization of gene expression measurements across platforms.
Winston Patrick Kuo, Tor-Kristian Jenssen, Atul J. Butte, Lucila Ohno-Machado, Isaac S. Kohane
Bioinform.3
2002 Comparing expression profiles of genes with similar promoter regions
abstract
MOTIVATION: Gene regulatory elements are often predicted by seeking common sequences in the promoter regions of genes that are clustered together based on their expression profiles. We consider the problem in the opposite direction: we seek to find the genes that have similar promoter regions and determine the extent to which these genes have similar expression profiles. RESULTS: We use the data sets from experiments on Saccharomyces cerevisiae. Our similarity measure for the promoter regions is based on the set of common mapped or putative transcription factor binding sites and other regulatory elements in the upstream region of the genes, as contained in the Saccharomyces cerevisiae Promoter Database. We pair up the genes with high similarity scores and compare their expression levels in time-course experiment data. We find that genes with similar promoter regions on the average have significantly higher correlation, but it can vary widely depending on the genes. This confirms that the presence of similar regulatory elements often does not correspond to similarity in expression profiles and indicates that finding transcription factor binding sites or other regulatory elements starting with the expression patterns may be limited in many cases. Regardless of the correlation, the degree to which the profiles agree under different experimental conditions can be examined to derive hypotheses concerning the role of common regulatory elements. Overall, we find that considering the relationship between the promoter regions and the expression profiles starting with the regulatory elements is a difficult but useful process that can provide valuable insights.
Peter J. Park, Atul J. Butte, Isaac S. Kohane
Bioinform.2
2001 Comparing the Similarity of Time-Series Gene Expression Using Signal Processing Metrics
Atul J. Butte, Ling Bao, Ben Y. Reis, Timothy W. Watkins, Isaac S. Kohane
J. Biomed. Informatics1
2001 Extracting Knowledge from Dynamics in Gene Expression
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane
J. Biomed. Informatics2
2001 Reply
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane
J. Biomed. Informatics2
2000 Enrolling patients into clinical trials faster using RealTime Recuiting
Atul J. Butte, David A. Weinstein, Isaac S. Kohane
AMIA1
1999 Unsupervised knowledge discovery in medical databases using relevance networks
Atul J. Butte, Isaac S. Kohane
AMIA1