Isaac S. Kohane

dblp:k/IsaacSKohane · DBLP profile ↗
← Back
143ranked-venue papers
11as first author
11since 2021 · last 2026
0000-0003-2192-5160ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 133 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Heterogenous effect of automated alerts on mortality
abstract
OBJECTIVE: To understand the heterogeneous treatment effects of electronic alerts for acute kidney injury (AKI). MATERIALS AND METHODS: Secondary analysis of individual patient data from 3 randomized controlled trials. Our outcome measure was 14-day all-cause mortality. Data from the ELAIA-1 trial were used to predict the individualized effect of alerts on mortality based on patients' phenotype. Results were internally validated on a holdout dataset and externally validated using data from 2 additional trials: UPenn and ELAIA-2. We used machine learning-based methods and performed a meta-analysis on individual patient data to identify patient subgroups whose risk of mortality was associated with alerts. In addition, provider actions following alerts were examined to explain how alerts impacted patient mortality. RESULTS: Compared to patients who were predicted to be harmed by an alert, patients predicted to benefit had a lower risk of death in both the internal validation cohort (n = 1809 patients; Pinteraction = .045) and both external validation cohorts (n = 7453 patients; Pinteraction < .0001). In external cohorts, 43 deaths may have been preventable if alerts were restricted to likely beneficiaries. Machine-learning based meta-analysis identified reduced mortality with alerts among patients with higher blood pressures (BP) and lower predicted risk, but increased mortality in non-urban and non-teaching hospitals. Provider responses to alerts differed across subgroups. DISCUSSION: Our findings indicate substantial heterogeneity in the effects of AKI alerts on patient mortality. Tailoring alert delivery based on predicted benefit may mitigate harm and enhance clinical outcomes. CONCLUSION: Individualizing automated alerts may reduce all-cause mortality. A prospective trial of individualized alerts is needed to confirm these results. TRIAL REGISTRATION: https://clinicaltrials.gov/ct2/show/NCT02753751 and https://clinicaltrials.gov/ct2/show/NCT02771977.
Benjamin D. Wissel, Zana Percy, Tanner J. Zachem, Brett K. Beaulieu-Jones, Isaac S. Kohane, Stuart L. Goldstein, Emrah Gecili, Judith W. Dexheimer
J. Am. Medical Informatics Assoc.5
2025 A standards-based approach to digital health research: implementing the people heart study
abstract
OBJECTIVE: To assess whether HL7 Fast Healthcare Interoperability Resources (FHIR) can underpin a fully standards-based, end-to-end digital research architecture, demonstrate it in a live study, and quantify its benefits for interoperability and development efficiency. MATERIALS AND METHODS: We designed a generalizable standards-based architecture to accelerate digital health research relying on FHIR as the sole transactional model throughout a participant research lifecycle starting from API-based study discovery to results. It was instantiated for People Heart Study, a real-world digital health cardiovascular-risk assessment study with its protocol transformed into FHIR resources (eligibility, consent, tasks, and results). Evaluation examined workflow coverage, validator conformance across independent servers, and points requiring custom extensions or app logic. RESULTS: The architecture was implemented using cloud managed FHIR stores including an illustrative public research discovery API for first-/third-party apps. A participant-facing iOS app was published on the App Store. Our evaluation reveals that 6 of 10 research app workflows could be executed entirely from FHIR artifacts; 2 were partially standards-driven and 2 remained limited requiring custom development. All FHIR resources passed structural, semantic validation with minimal custom extension usage and terminology integrity issues. DISCUSSION: Our approach addresses persistent challenges in digital health research by enhancing data interoperability, minimizing redundant development, and supporting the full research lifecycle. The architecture aligns with national priorities and complements healthcare standardization efforts. CONCLUSION: By leveraging FHIR, our architecture enables generalizability, interoperability, and reuse across diverse digital health research contexts, transforming study design into data modeling rather than software development, and fostering a more inclusive and agile digital health ecosystem.
Raheel Sayeed, David A. Kreda, Joshua C. Mandel, Bryan Larson, William J. Gordon, Kenneth D. Mandl, Isaac S. Kohane
J. Am. Medical Informatics Assoc.7
2023 Event Stream GPT: A Data Pre-processing and Modeling Library for Generative, Pre-trained Transformers over Continuous-time Sequences of Complex Events
abstract
Generative, pre-trained transformers (GPTs, a type of "Foundation Models") have reshaped natural language processing (NLP) through their versatility in diverse downstream tasks. However, their potential extends far beyond NLP. This paper provides a software utility to help realize this potential, extending the applicability of GPTs to continuous-time sequences of complex events with internal dependencies, such as medical record datasets. Despite their potential, the adoption of foundation models in these domains has been hampered by the lack of suitable tools for model construction and evaluation. To bridge this gap, we introduce Event Stream GPT (ESGPT), an open-source library designed to streamline the end-to-end process for building GPTs for continuous-time event sequences. ESGPT allows users to (1) build flexible, foundation-model scale input datasets by specifying only a minimal configuration file, (2) leverage a Hugging Face compatible modeling API for GPTs over this modality that incorporates intra-event causal dependency structures and autoregressive generation capabilities, and (3) evaluate models via standardized processes that can assess few and even zero-shot performance of pre-trained models on user-specified fine-tuning tasks.
Matthew B. A. McDermott, Bret Nestor, Peniel N. Argaw, Isaac S. Kohane
NeurIPS4
2022 Multi-PheWAS intersection approach to identify sex differences across comorbidities in 59 140 pediatric patients with autism spectrum disorder
abstract
OBJECTIVE: To identify differences related to sex and define autism spectrum disorder (ASD) comorbidities female-enriched through a comprehensive multi-PheWAS intersection approach on big, real-world data. Although sex difference is a consistent and recognized feature of ASD, additional clinical correlates could help to identify potential disease subgroups, based on sex and age. MATERIALS AND METHODS: We performed a systematic comorbidity analysis on 1860 groups of comorbidities exploring all spectrum of known disease, in 59 140 individuals (11 440 females) with ASD from 4 age groups. We explored ASD sex differences in 2 independent real-world datasets, across all potential comorbidities by comparing (1) females with ASD vs males with ASD and (2) females with ASD vs females without ASD. RESULTS: We identified 27 different comorbidities that appeared significantly more frequently in females with ASD. The comorbidities were mostly neurological (eg, epilepsy, odds ratio [OR] > 1.8, 3-18 years of age), congenital (eg, chromosomal anomalies, OR > 2, 3-18 years of age), and mental disorders (eg, intellectual disability, OR > 1.7, 6-18 years of age). Novel comorbidities included endocrine metabolic diseases (eg, failure to thrive, OR = 2.5, ages 0-2), digestive disorders (gastroesophageal reflux disease: OR = 1.7, 6-11 years of age; and constipation: OR > 1.6, 3-11 years of age), and sense organs (strabismus: OR > 1.8, 3-18 years of age). DISCUSSION: A multi-PheWAS intersection approach on real-world data as presented in this study uniquely contributes to the growing body of research regarding sex-based comorbidity analysis in ASD population. CONCLUSIONS: Our findings provide insights into female-enriched ASD comorbidities that are potentially important in diagnosis, as well as the identification of distinct comorbidity patterns influencing anticipatory treatment or referrals. The code is publicly available (https://github.com/hms-dbmi/sexDifferenceInASD).
Alba Gutiérrez-Sacristán, Carlos Sáez 0001, Carlos De Niz, Niloofar Jalali, Thomas N. Desain, Ranjay Kumar, Joany M. Zachariasse, Kathe P. Fox, Nathan P. Palmer, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.10
2022 Analytics to monitor local impact of the Protecting Access to Medicare Act's imaging clinical decision support requirements
abstract
OBJECTIVE: This study aimed is to: (1) extend the Integrating the Biology and the Bedside (i2b2) data and application models to include medical imaging appropriate use criteria, enabling it to serve as a platform to monitor local impact of the Protecting Access to Medicare Act's (PAMA) imaging clinical decision support (CDS) requirements, and (2) validate the i2b2 extension using data from the Medicare Imaging Demonstration (MID) CDS implementation. MATERIALS AND METHODS: This study provided a reference implementation and assessed its validity and reliability using data from the MID, the federal government's predecessor to PAMA's imaging CDS program. The Star Schema was extended to describe the interactions of imaging ordering providers with the CDS. New ontologies were added to enable mapping medical imaging appropriateness data to i2b2 schema. z-Ratio for testing the significance of the difference between 2 independent proportions was utilized. RESULTS: The reference implementation used 26 327 orders for imaging examinations which were persisted to the modified i2b2 schema. As an illustration of the analytical capabilities of the Web Client, we report that 331/1192 or 28.1% of imaging orders were deemed appropriate by the CDS system at the end of the intervention period (September 2013), an increase from 162/1223 or 13.2% for the first month of the baseline period, December 2011 (P = .0212), consistent with previous studies. CONCLUSIONS: The i2b2 platform can be extended to monitor local impact of PAMA's appropriateness of imaging ordering CDS requirements.
Vladimir I. Valtchinov, Shawn N. Murphy, Ronilda C. Lacson, Nikolay Ikonomov, Bingxue K. Zhai, Katherine P. Andriole, Justin F. Rousseau, Dick Hanson, Isaac S. Kohane, Ramin Khorasani
J. Am. Medical Informatics Assoc.9
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics55
2021 Patient-led data sharing for clinical bioinformatics research: USCDI and beyond
abstract
The 21st Century Cures Act, passed in 2016, and the Final Rules it called for create a roadmap for enabling patient access to their electronic health information. The set of data to be made available, as determined by the Office of the National Coordinator for Health IT through the US Core Data for Interoperability expansion process, will impact the value creation of this improved data liquidity. In this commentary, we look at the potential for significant value creation from USCDI in the context of clinical bioinformatics research and advocate for the research community's involvement in the USCDI process to propel this value creation forward. We also describe 1 mechanism-using existing required APIs for full data export capabilities-that could pragmatically enable this value creation at minimal additional technical lift beyond the current regulatory requirements.
William J. Gordon, Daniel Gottlieb 0001, David A. Kreda, Joshua C. Mandel, Kenneth D. Mandl, Isaac S. Kohane
J. Am. Medical Informatics Assoc.6
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.39
2021 Use of electronic health records to support a public health response to the COVID-19 pandemic in the United States: a perspective from 15 academic medical centers
abstract
Our goal is to summarize the collective experience of 15 organizations in dealing with uncoordinated efforts that result in unnecessary delays in understanding, predicting, preparing for, containing, and mitigating the COVID-19 pandemic in the US. Response efforts involve the collection and analysis of data corresponding to healthcare organizations, public health departments, socioeconomic indicators, as well as additional signals collected directly from individuals and communities. We focused on electronic health record (EHR) data, since EHRs can be leveraged and scaled to improve clinical care, research, and to inform public health decision-making. We outline the current challenges in the data ecosystem and the technology infrastructure that are relevant to COVID-19, as witnessed in our 15 institutions. The infrastructure includes registries and clinical data networks to support population-level analyses. We propose a specific set of strategic next steps to increase interoperability, overall organization, and efficiencies.
Subha Madhavan, Lisa Bastarache, Jeffrey S. Brown, Atul J. Butte, David A. Dorr, Peter J. Embí, Charles P. Friedman, Kevin B. Johnson, Jason H. Moore, Isaac S. Kohane, Philip R. O. Payne, Jessica D. Tenenbaum, Mark G. Weiner, Adam B. Wilcox, Lucila Ohno-Machado
J. Am. Medical Informatics Assoc.10
2021 Finding commonalities in rare diseases through the undiagnosed diseases network
abstract
OBJECTIVE: When studying any specific rare disease, heterogeneity and scarcity of affected individuals has historically hindered investigators from discerning on what to focus to understand and diagnose a disease. New nongenomic methodologies must be developed that identify similarities in seemingly dissimilar conditions. MATERIALS AND METHODS: This observational study analyzes 1042 patients from the Undiagnosed Diseases Network (2015-2019), a multicenter, nationwide research study using phenotypic data annotated by specialized staff using Human Phenotype Ontology terms. We used Louvain community detection to cluster patients linked by Jaccard pairwise similarity and 2 support vector classifier to assign new cases. We further validated the clusters' most representative comorbidities using a national claims database (67 million patients). RESULTS: Patients were divided into 2 groups: those with symptom onset before 18 years of age (n = 810) and at 18 years of age or older (n = 232) (average symptom onset age: 10 [interquartile range, 0-14] years). For 810 pediatric patients, we identified 4 statistically significant clusters. Two clusters were characterized by growth disorders, and developmental delay enriched for hypotonia presented a higher likelihood of diagnosis. Support vector classifier showed 0.89 balanced accuracy (0.83 for Human Phenotype Ontology terms only) on test data. DISCUSSIONS: To set the framework for future discovery, we chose as our endpoint the successful grouping of patients by phenotypic similarity and provide a classification tool to assign new patients to those clusters. CONCLUSION: This study shows that despite the scarcity and heterogeneity of patients, we can still find commonalities that can potentially be harnessed to uncover new insights and targets for therapy.
Josephine Yates, Alba Gutiérrez-Sacristán, Vianney Jouhet, Kimberly Leblanc, Cecilia Esteves, Thomas N. Desain, Nick Benik, Jason Stedman, Nathan P. Palmer, Guillaume Mellon, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.11
2021 ATLAS: an automated association test using probabilistically linked health records with application to genetic studies
abstract
OBJECTIVE: Large amounts of health data are becoming available for biomedical research. Synthesizing information across databases may capture more comprehensive pictures of patient health and enable novel research studies. When no gold standard mappings between patient records are available, researchers may probabilistically link records from separate databases and analyze the linked data. However, previous linked data inference methods are constrained to certain linkage settings and exhibit low power. Here, we present ATLAS, an automated, flexible, and robust association testing algorithm for probabilistically linked data. MATERIALS AND METHODS: Missing variables are imputed at various thresholds using a weighted average method that propagates uncertainty from probabilistic linkage. Next, estimated effect sizes are obtained using a generalized linear model. ATLAS then conducts the threshold combination test by optimally combining P values obtained from data imputed at varying thresholds using Fisher's method and perturbation resampling. RESULTS: In simulations, ATLAS controls for type I error and exhibits high power compared to previous methods. In a real-world genetic association study, meta-analysis of ATLAS-enabled analyses on a linked cohort with analyses using an existing cohort yielded additional significant associations between rheumatoid arthritis genetic risk score and laboratory biomarkers. DISCUSSION: Weighted average imputation weathers false matches and increases contribution of true matches to mitigate linkage error-induced bias. The threshold combination test avoids arbitrarily choosing a threshold to rule a match, thus automating linked data-enabled analyses and preserving power. CONCLUSION: ATLAS promises to enable novel and powerful research studies using linked data to capitalize on all available data sources.
Harrison G. Zhang, Boris P. Hejblum, Griffin M. Weber, Nathan P. Palmer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Katherine P. Liao, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.9
2020 Deciphering Serous Ovarian Carcinoma Histopathology and Platinum Response by Convolutional Neural Networks
Kun-Hsing Yu, Vincent Hu, Ursula Matulonis, George Mutter, Jeffrey Golden, Isaac S. Kohane
AMIA7
2020 Meta-analysis of Caenorhabditis elegans single-cell developmental data reveals multi-frequency oscillation in gene activation
abstract
MOTIVATION: The advent of in vivo automated techniques for single-cell lineaging, sequencing and analysis of gene expression has begun to dramatically increase our understanding of organismal development. We applied novel meta-analysis and visualization techniques to the EPIC single-cell-resolution developmental gene expression dataset for Caenorhabditis elegans from Bao, Murray, Waterston et al. to gain insights into regulatory mechanisms governing the timing of development. RESULTS: Our meta-analysis of the EPIC dataset revealed that a simple linear combination of the expression levels of the developmental genes is strongly correlated with the developmental age of the organism, irrespective of the cell division rate of different cell lineages. We uncovered a pattern of collective sinusoidal oscillation in gene activation, in multiple dominant frequencies and in multiple orthogonal axes of gene expression, pointing to the existence of a coordinated, multi-frequency global timing mechanism. We developed a novel method based on Fisher's Discriminant Analysis to identify gene expression weightings that maximally separate traits of interest, and found that remarkably, simple linear gene expression weightings are capable of producing sinusoidal oscillations of any frequency and phase, adding to the growing body of evidence that oscillatory mechanisms likely play an important role in the timing of development. We cross-linked EPIC with gene ontology and anatomy ontology terms, employing Fisher's Discriminant Analysis methods to identify previously unknown positive and negative genetic contributions to developmental processes and cell phenotypes. This meta-analysis demonstrates new evidence for direct linear and/or sinusoidal mechanisms regulating the timing of development. We uncovered a number of previously unknown positive and negative correlations between developmental genes and developmental processes or cell phenotypes. Our results highlight both the continued relevance of the EPIC technique, and the value of meta-analysis of previously published results. The presented analysis and visualization techniques are broadly applicable across developmental and systems biology. AVAILABILITY AND IMPLEMENTATION: Analysis software available upon request. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Luke A. D. Hutchison, Bonnie Berger, Isaac S. Kohane
Bioinform.3
2020 Scalability and cost-effectiveness analysis of whole genome-wide association studies on Google Cloud Platform and Amazon Web Services
abstract
OBJECTIVE: Advancements in human genomics have generated a surge of available data, fueling the growth and accessibility of databases for more comprehensive, in-depth genetic studies. METHODS: We provide a straightforward and innovative methodology to optimize cloud configuration in order to conduct genome-wide association studies. We utilized Spark clusters on both Google Cloud Platform and Amazon Web Services, as well as Hail (http://doi.org/10.5281/zenodo.2646680) for analysis and exploration of genomic variants dataset. RESULTS: Comparative evaluation of numerous cloud-based cluster configurations demonstrate a successful and unprecedented compromise between speed and cost for performing genome-wide association studies on 4 distinct whole-genome sequencing datasets. Results are consistent across the 2 cloud providers and could be highly useful for accelerating research in genetics. CONCLUSIONS: We present a timely piece for one of the most frequently asked questions when moving to the cloud: what is the trade-off between speed and cost?
Inès Krissaane, Carlos De Niz, Alba Gutiérrez-Sacristán, Gabor Korodi, Nneka Ede, Ranjay Kumar, Jessica Lyons, Arjun K. Manrai, Chirag J. Patel, Isaac S. Kohane, Paul Avillach
J. Am. Medical Informatics Assoc.10
2020 An interactive online dashboard for tracking COVID-19 in U.S. counties, cities, and states in real time
abstract
OBJECTIVE: The study sought to create an online resource that informs the public of coronavirus disease 2019 (COVID-19) outbreaks in their area. MATERIALS AND METHODS: This R Shiny application aggregates data from multiple resources that track COVID-19 and visualizes them through an interactive, online dashboard. RESULTS: The Web resource, called the COVID-19 Watcher, can be accessed online (https://covid19watcher.research.cchmc.org/). It displays COVID-19 data from every county and 188 metropolitan areas in the United States. Features include rankings of the worst-affected areas and auto-generating plots that depict temporal changes in testing capacity, cases, and deaths. DISCUSSION: The Centers for Disease Control and Prevention does not publish COVID-19 data for local municipalities, so it is critical that academic resources fill this void so the public can stay informed. The data used have limitations and likely underestimate the scale of the outbreak. CONCLUSIONS: The COVID-19 Watcher can provide the public with real-time updates of outbreaks in their area.
Benjamin D. Wissel, P. J. Van Camp, Michal Kouril, Chad Weis, Tracy A. Glauser, Peter S. White, Isaac S. Kohane, Judith W. Dexheimer
J. Am. Medical Informatics Assoc.7
2020 Classifying non-small cell lung cancer types and transcriptomic subtypes using convolutional neural networks
abstract
OBJECTIVE: Non-small cell lung cancer is a leading cause of cancer death worldwide, and histopathological evaluation plays the primary role in its diagnosis. However, the morphological patterns associated with the molecular subtypes have not been systematically studied. To bridge this gap, we developed a quantitative histopathology analytic framework to identify the types and gene expression subtypes of non-small cell lung cancer objectively. MATERIALS AND METHODS: We processed whole-slide histopathology images of lung adenocarcinoma (n = 427) and lung squamous cell carcinoma patients (n = 457) in the Cancer Genome Atlas. We built convolutional neural networks to classify histopathology images, evaluated their performance by the areas under the receiver-operating characteristic curves (AUCs), and validated the results in an independent cohort (n = 125). RESULTS: To establish neural networks for quantitative image analyses, we first built convolutional neural network models to identify tumor regions from adjacent dense benign tissues (AUCs > 0.935) and recapitulated expert pathologists' diagnosis (AUCs > 0.877), with the results validated in an independent cohort (AUCs = 0.726-0.864). We further demonstrated that quantitative histopathology morphology features identified the major transcriptomic subtypes of both adenocarcinoma and squamous cell carcinoma (P < .01). DISCUSSION: Our study is the first to classify the transcriptomic subtypes of non-small cell lung cancer using fully automated machine learning methods. Our approach does not rely on prior pathology knowledge and can discover novel clinically relevant histopathology patterns objectively. The developed procedure is generalizable to other tumor types or diseases.
Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane
J. Am. Medical Informatics Assoc.7
2020 Deep Learning Benchmarks on L1000 Gene Expression Data
abstract
Gene expression data can offer deep, physiological insights beyond the static coding of the genome alone. We believe that realizing this potential requires specialized, high-capacity machine learning methods capable of using underlying biological structure, but the development of such models is hampered by the lack of published benchmark tasks and well characterized baselines. In this work, we establish such benchmarks and baselines by profiling many classifiers against biologically motivated tasks on two curated views of a large, public gene expression dataset (the LINCS corpus) and one privately produced dataset. We provide these two curated views of the public LINCS dataset and our benchmark tasks to enable direct comparisons to future methodological work and help spur deep learning method development on this modality. In addition to profiling a battery of traditional classifiers, including linear models, random forests, decision trees, K nearest neighbor (KNN) classifiers, and feed-forward artificial neural networks (FF-ANNs), we also test a method novel to this data modality: graph convolugtional neural networks (GCNNs), which allow us to incorporate prior biological domain knowledge. We find that GCNNs can be highly performant, with large datasets, whereas FF-ANNs consistently perform well. Non-neural classifiers are dominated by linear models and KNN classifiers.
Matthew B. A. McDermott, Wen-Ning Zhao, Steven Sheridan, Peter Szolovits, Isaac S. Kohane, Stephen J. Haggarty, Roy H. Perlis
IEEE ACM Trans. Comput. Biol. Bioinform.6
2019 Learning to Estimate Nutrition Facts from Food Descriptions
Hadi Amiri, Andrew L. Beam, Isaac S. Kohane
AMIA3
2019 Classifying Non-Small Cell Lung Cancer Histopathology Types and Transcriptomic Subtypes using Convolutional Neural Networks
Kun-Hsing Yu, Gerald J. Berry, Christopher Ré, Russ B. Altman, Michael Snyder 0001, Isaac S. Kohane
AMIA7
2019 Batch correction evaluation framework using a-priori gene-gene associations: applied to the GTEx dataset
abstract
BACKGROUND: Correcting a heterogeneous dataset that presents artefacts from several confounders is often an essential bioinformatics task. Attempting to remove these batch effects will result in some biologically meaningful signals being lost. Thus, a central challenge is assessing if the removal of unwanted technical variation harms the biological signal that is of interest to the researcher. RESULTS: We describe a novel framework, B-CeF, to evaluate the effectiveness of batch correction methods and their tendency toward over or under correction. The approach is based on comparing co-expression of adjusted gene-gene pairs to a-priori knowledge of highly confident gene-gene associations based on thousands of unrelated experiments derived from an external reference. Our framework includes three steps: (1) data adjustment with the desired methods (2) calculating gene-gene co-expression measurements for adjusted datasets (3) evaluating the performance of the co-expression measurements against a gold standard. Using the framework, we evaluated five batch correction methods applied to RNA-seq data of six representative tissue datasets derived from the GTEx project. CONCLUSIONS: Our framework enables the evaluation of batch correction methods to better preserve the original biological signal. We show that using a multiple linear regression model to correct for known confounders outperforms factor analysis-based methods that estimate hidden confounders. The code is publicly available as an R package.
Judith Somekh, Shai S. Shen-Orr, Isaac S. Kohane
BMC Bioinform.3
2019 High-throughput multimodal automated phenotyping (MAP) with application to PheWAS
abstract
OBJECTIVE: Electronic health records linked with biorepositories are a powerful platform for translational studies. A major bottleneck exists in the ability to phenotype patients accurately and efficiently. The objective of this study was to develop an automated high-throughput phenotyping method integrating International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP). MATERIALS AND METHODS: We developed a mapping method for automatically identifying relevant ICD and NLP concepts for a specific phenotype leveraging the Unified Medical Language System. Along with health care utilization, aggregated ICD and NLP counts were jointly analyzed by fitting an ensemble of latent mixture models. The multimodal automated phenotyping (MAP) algorithm yields a predicted probability of phenotype for each patient and a threshold for classifying participants with phenotype yes/no. The algorithm was validated using labeled data for 16 phenotypes from a biorepository and further tested in an independent cohort phenome-wide association studies (PheWAS) for 2 single nucleotide polymorphisms with known associations. RESULTS: The MAP algorithm achieved higher or similar AUC and F-scores compared to the ICD code across all 16 phenotypes. The features assembled via the automated approach had comparable accuracy to those assembled via manual curation (AUCMAP 0.943, AUCmanual 0.941). The PheWAS results suggest that the MAP approach detected previously validated associations with higher power when compared to the standard PheWAS method based on ICD codes. CONCLUSION: The MAP approach increased the accuracy of phenotype definition while maintaining scalability, thereby facilitating use in studies requiring large-scale phenotyping, such as PheWAS.
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Yuk-Lam Ho, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Christopher J. O'Donnell, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002
J. Am. Medical Informatics Assoc.18
2019 Feature extraction for phenotyping from semantic and knowledge resources
Wenxin Ning, Stephanie Chan 0002, Andrew L. Beam, Ming Yu 0003, Alon Geva, Katherine P. Liao, Mary Mullen, Kenneth D. Mandl, Isaac S. Kohane, Tianxi Cai, Sheng Yu 0002
J. Biomed. Informatics9
2019 Automated grouping of medical codes via multiview banded spectral clustering
abstract
OBJECTIVE: With its increasingly widespread adoption, electronic health records (EHR) have enabled phenotypic information extraction at an unprecedented granularity and scale. However, often a medical concept (e.g. diagnosis, prescription, symptom) is described in various synonyms across different EHR systems, hindering data integration for signal enhancement and complicating dimensionality reduction for knowledge discovery. Despite existing ontologies and hierarchies, tremendous human effort is needed for curation and maintenance - a process that is both unscalable and susceptible to subjective biases. This paper aims to develop a data-driven approach to automate grouping medical terms into clinically relevant concepts by combining multiple up-to-date data sources in an unbiased manner. METHODS: We present a novel data-driven grouping approach - multi-view banded spectral clustering (mvBSC) combining summary data from multiple healthcare systems. The proposed method consists of a banding step that leverages the prior knowledge from the existing coding hierarchy, and a combining step that performs spectral clustering on an optimally weighted matrix. RESULTS: -measure, and were found to consistently exhibit great similarity to the existing manual grouping counterpart. The resulting ICD groupings also enjoy comparable interpretability and are well aligned with the current ICD hierarchy. CONCLUSION: The proposed approach, by systematically leveraging multiple data sources, is able to overcome bias while maximizing consensus to achieve generalizability. It has the advantage of being efficient, scalable, and adaptive to the evolving human knowledge reflected in the data, showing a significant step toward automating medical knowledge integration.
Luwan Zhang, Tianrun A. Cai, Yuri Ahuja, Zeling He, Yuk-Lam Ho, Andrew L. Beam, Kelly Cho, Robert J. Carroll, Joshua C. Denny, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai
J. Biomed. Informatics11
2018 High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai
AMIA16
2018 Rcupcake: an R package for querying and analyzing biomedical data through the BD2K PIC-SURE RESTful API
abstract
Motivation: In the era of big data and precision medicine, the number of databases containing clinical, environmental, self-reported and biochemical variables is increasing exponentially. Enabling the experts to focus on their research questions rather than on computational data management, access and analysis is one of the most significant challenges nowadays. Results: We present Rcupcake, an R package that contains a variety of functions for leveraging different databases through the BD2K PIC-SURE RESTful API and facilitating its query, analysis and interpretation. The package offers a variety of analysis and visualization tools, including the study of the phenotype co-occurrence and prevalence, according to multiple layers of data, such as phenome, exposome or genome. Availability and implementation: The package is implemented in R and is available under Mozilla v2 license from GitHub (https://github.com/hms-dbmi/Rcupcake). Two reproducible case studies are also available (https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/SSCcaseStudy_v01.ipynb, https://github.com/hms-dbmi/Rcupcake-case-studies/blob/master/NHANEScaseStudy_v01.ipynb). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Alba Gutiérrez-Sacristán, Romain Guedj, Gabor Korodi, Jason Stedman, Laura Inés Furlong, Chirag J. Patel, Isaac S. Kohane, Paul Avillach
Bioinform.7
2018 PheProb: probabilistic phenotyping using diagnosis codes to improve power for genetic association studies
abstract
Objective: Standard approaches for large scale phenotypic screens using electronic health record (EHR) data apply thresholds, such as ≥2 diagnosis codes, to define subjects as having a phenotype. However, the variation in the accuracy of diagnosis codes can impair the power of such screens. Our objective was to develop and evaluate an approach which converts diagnosis codes into a probability of a phenotype (PheProb). We hypothesized that this alternate approach for defining phenotypes would improve power for genetic association studies. Methods: The PheProb approach employs unsupervised clustering to separate patients into 2 groups based on diagnosis codes. Subjects are assigned a probability of having the phenotype based on the number of diagnosis codes. This approach was developed using simulated EHR data and tested in a real world EHR cohort. In the latter, we tested the association between low density lipoprotein cholesterol (LDL-C) genetic risk alleles known for association with hyperlipidemia and hyperlipidemia codes (ICD-9 272.x). PheProb and thresholding approaches were compared. Results: Among n = 1462 subjects in the real world EHR cohort, the threshold-based p-values for association between the genetic risk score (GRS) and hyperlipidemia were 0.126 (≥1 code), 0.123 (≥2 codes), and 0.142 (≥3 codes). The PheProb approach produced the expected significant association between the GRS and hyperlipidemia: p = .001. Conclusions: PheProb improves statistical power for association studies relative to standard thresholding approaches by leveraging information about the phenotype in the billing code counts. The PheProb approach has direct applications where efficient approaches are required, such as in Phenome-Wide Association Studies.
Jennifer A. Sinnott, Fiona Cai, Sheng Yu 0002, Boris P. Hejblum, Chuan Hong, Isaac S. Kohane, Katherine P. Liao
J. Am. Medical Informatics Assoc.6
2018 Enabling phenotypic big data with PheNorm
abstract
Objective: Electronic health record (EHR)-based phenotyping infers whether a patient has a disease based on the information in his or her EHR. A human-annotated training set with gold-standard disease status labels is usually required to build an algorithm for phenotyping based on a set of predictive features. The time intensiveness of annotation and feature curation severely limits the ability to achieve high-throughput phenotyping. While previous studies have successfully automated feature curation, annotation remains a major bottleneck. In this paper, we present PheNorm, a phenotyping algorithm that does not require expert-labeled samples for training. Methods: The most predictive features, such as the number of International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes or mentions of the target phenotype, are normalized to resemble a normal mixture distribution with high area under the receiver operating curve (AUC) for prediction. The transformed features are then denoised and combined into a score for accurate disease classification. Results: We validated the accuracy of PheNorm with 4 phenotypes: coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis. The AUCs of the PheNorm score reached 0.90, 0.94, 0.95, and 0.94 for the 4 phenotypes, respectively, which were comparable to the accuracy of supervised algorithms trained with sample sizes of 100-300, with no statistically significant difference. Conclusion: The accuracy of the PheNorm algorithms is on par with algorithms trained with annotated samples. PheNorm fully automates the generation of accurate phenotyping algorithms and demonstrates the capacity for EHR-driven annotations to scale to the next level - phenotypic big data.
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai
J. Am. Medical Informatics Assoc.10
2017 High-throughput Phenotyping via Denoised Normal Mixture Transformation
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
AMIA11
2017 Subtype Discovery for Autism Spectrum Disorder by Comorbidity Analysis
Kun-Hsing Yu, Nathan P. Palmer, Isaac S. Kohane
AMIA3
2017 Surrogate-assisted feature extraction for high-throughput phenotyping
abstract
OBJECTIVE: Phenotyping algorithms are capable of accurately identifying patients with specific phenotypes from within electronic medical records systems. However, developing phenotyping algorithms in a scalable way remains a challenge due to the extensive human resources required. This paper introduces a high-throughput unsupervised feature selection method, which improves the robustness and scalability of electronic medical record phenotyping without compromising its accuracy. METHODS: The proposed Surrogate-Assisted Feature Extraction (SAFE) method selects candidate features from a pool of comprehensive medical concepts found in publicly available knowledge sources. The target phenotype's International Classification of Diseases, Ninth Revision and natural language processing counts, acting as noisy surrogates to the gold-standard labels, are used to create silver-standard labels. Candidate features highly predictive of the silver-standard labels are selected as the final features. RESULTS: Algorithms were trained to identify patients with coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis using various numbers of labels to compare the performance of features selected by SAFE, a previously published automated feature extraction for phenotyping procedure, and domain experts. The out-of-sample area under the receiver operating characteristic curve and F -score from SAFE algorithms were remarkably higher than those from the other two, especially at small label sizes. CONCLUSION: SAFE advances high-throughput phenotyping methods by automatically selecting a succinct set of informative features for algorithm training, which in turn reduces overfitting and the needed number of gold-standard labels. SAFE also potentially identifies important features missed by automated feature extraction for phenotyping or experts.
Sheng Yu 0002, Abhishek Chakrabortty, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.10
2016 The Spectrum of Insomnia-Associated Comorbidities in an Electronic Medical Records Cohort
Uri Kartoun, Andrew L. Beam, Jennifer Pai, Arnaub Chatterjee, Timothy P. Fitzgerald, Isaac S. Kohane, Stanley Y. Shaw
AMIA6
2016 Characterizing the Fever Effect in Autism Spectrum Disorder
Efrat Muller, Alal Eran, Denis Agniel, Isaac S. Kohane, Eitan Bachmat
AMIA4
2016 ksRepo: a generalized platform for computational drug repositioning
abstract
BACKGROUND: Repositioning approved drug and small molecules in novel therapeutic areas is of key interest to the pharmaceutical industry. A number of promising computational techniques have been developed to aid in repositioning, however, the majority of available methodologies require highly specific data inputs that preclude the use of many datasets and databases. There is a clear unmet need for a generalized methodology that enables the integration of multiple types of both gene expression data and database schema. RESULTS: ksRepo eliminates the need for a single microarray platform as input and allows for the use of a variety of drug and chemical exposure databases. We tested ksRepo's performance on a set of five prostate cancer datasets using the Comparative Toxicogenomics Database (CTD) as our database of gene-compound interactions. ksRepo successfully predicted significance for five frontline prostate cancer therapies, representing a significant enrichment from over 7000 CTD compounds, and achieved specificity similar to other repositioning methods. CONCLUSIONS: We present ksRepo, which enables investigators to use any data inputs for computational drug repositioning. ksRepo is implemented in a series of four functions in the R statistical environment under a BSD3 license. Source code is freely available at http://github.com/adam-sam-brown/ksRepo. A vignette is provided to aid users in performing ksRepo analysis.
Adam S. Brown, Sek Won Kong, Isaac S. Kohane, Chirag J. Patel
BMC Bioinform.3
2016 SMART on FHIR: a standards-based, interoperable apps platform for electronic health records
abstract
OBJECTIVE: In early 2010, Harvard Medical School and Boston Children's Hospital began an interoperability project with the distinctive goal of developing a platform to enable medical applications to be written once and run unmodified across different healthcare IT systems. The project was called Substitutable Medical Applications and Reusable Technologies (SMART). METHODS: We adopted contemporary web standards for application programming interface transport, authorization, and user interface, and standard medical terminologies for coded data. In our initial design, we created our own openly licensed clinical data models to enforce consistency and simplicity. During the second half of 2013, we updated SMART to take advantage of the clinical data models and the application-programming interface described in a new, openly licensed Health Level Seven draft standard called Fast Health Interoperability Resources (FHIR). Signaling our adoption of the emerging FHIR standard, we called the new platform SMART on FHIR. RESULTS: We introduced the SMART on FHIR platform with a demonstration that included several commercial healthcare IT vendors and app developers showcasing prototypes at the Health Information Management Systems Society conference in February 2014. This established the feasibility of SMART on FHIR, while highlighting the need for commonly accepted pragmatic constraints on the base FHIR specification. CONCLUSION: In this paper, we describe the creation of SMART on FHIR, relate the experience of the vendors and developers who built SMART on FHIR prototypes, and discuss some challenges in going from early industry prototyping to industry-wide production use.
Joshua C. Mandel, David A. Kreda, Kenneth D. Mandl, Isaac S. Kohane, Rachel Badovinac Ramoni
J. Am. Medical Informatics Assoc.4
2016 SMART precision cancer medicine: a FHIR-based app to provide genomic information at the point of care
abstract
BACKGROUND: Precision cancer medicine (PCM) will require ready access to genomic data within the clinical workflow and tools to assist clinical interpretation and enable decisions. Since most electronic health record (EHR) systems do not yet provide such functionality, we developed an EHR-agnostic, clinico-genomic mobile app to demonstrate several features that will be needed for point-of-care conversations. METHODS: Our prototype, called Substitutable Medical Applications and Reusable Technology (SMART)® PCM, visualizes genomic information in real time, comparing a patient's diagnosis-specific somatic gene mutations detected by PCR-based hotspot testing to a population-level set of comparable data. The initial prototype works for patient specimens with 0 or 1 detected mutation. Genomics extensions were created for the Health Level Seven® Fast Healthcare Interoperability Resources (FHIR)® standard; otherwise, the prototype is a normal SMART on FHIR app. RESULTS: The PCM prototype can rapidly present a visualization that compares a patient's somatic genomic alterations against a distribution built from more than 3000 patients, along with context-specific links to external knowledge bases. Initial evaluation by oncologists provided important feedback about the prototype's strengths and weaknesses. We added several requested enhancements and successfully demonstrated the app at the inaugural American Society of Clinical Oncology Interoperability Demonstration; we have also begun to expand visualization capabilities to include cancer specimens with multiple mutations. DISCUSSION: PCM is open-source software for clinicians to present the individual patient within the population-level spectrum of cancer somatic mutations. The app can be implemented on any SMART on FHIR-enabled EHRs, and future versions of PCM should be able to evolve in parallel with external knowledge bases.
Jeremy L. Warner, Matthew J. Rioth, Kenneth D. Mandl, Joshua C. Mandel, David A. Kreda, Isaac S. Kohane, Daniel Carbone, Ross Oreto, Lucy Wang, Shilin Zhu, Heming Yao, Gil Alterovitz
J. Am. Medical Informatics Assoc.6
2016 A model-driven methodology for exploring complex disease comorbidities applied to autism spectrum disorder and inflammatory bowel disease
Judith Somekh, Mor Peleg, Alal Eran, Itay Koren, Ariel Feiglin, Alik Demishtein, Ruth Shiloh, Monika Heiner, Sek Won Kong, Zvulun Elazar, Isaac S. Kohane
J. Biomed. Informatics11
2015 Demonstrating the Advantages of Applying Data Mining Techniques on Time-Dependent Electronic Medical Records
Uri Kartoun, Vishesh Kumar, Su-Chun Cheng, Sheng Yu 0002, Katherine P. Liao, Elizabeth W. Karlson, Ashwin N. Ananthakrishnan, Zongqi Xia, Vivian S. Gainer, Andrew Cagan, Guergana K. Savova, Pei J. Chen, Shawn N. Murphy, Susanne E. Churchill, Isaac S. Kohane, Peter Szolovits, Tianxi Cai, Stanley Y. Shaw
AMIA15
2015 SMART on FHIR Genomics: facilitating standardized clinico-genomic apps
abstract
BACKGROUND: Supporting clinical decision support for personalized medicine will require linking genome and phenome variants to a patient's electronic health record (EHR), at times on a vast scale. Clinico-genomic data standards will be needed to unify how genomic variant data are accessed from different sequencing systems. METHODS: A specification for the basis of a clinic-genomic standard, building upon the current Health Level Seven International Fast Healthcare Interoperability Resources (FHIR®) standard, was developed. An FHIR application protocol interface (API) layer was attached to proprietary sequencing platforms and EHRs in order to expose gene variant data for presentation to the end-user. Three representative apps based on the SMART platform were built to test end-to-end feasibility, including integration of genomic and clinical data. RESULTS: Successful design, deployment, and use of the API was demonstrated and adopted by HL7 Clinical Genomics Workgroup. Feasibility was shown through development of three apps by various types of users with background levels and locations. CONCLUSION: This prototyping work suggests that an entirely data (and web) standards-based approach could prove both effective and efficient for advancing personalized medicine.
Gil Alterovitz, Jeremy L. Warner, Peijin Zhang, Yishen Chen, Mollie Ullman-Cullere, David A. Kreda, Isaac S. Kohane
J. Am. Medical Informatics Assoc.7
2015 Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sources
abstract
OBJECTIVE: Analysis of narrative (text) data from electronic health records (EHRs) can improve population-scale phenotyping for clinical and genetic research. Currently, selection of text features for phenotyping algorithms is slow and laborious, requiring extensive and iterative involvement by domain experts. This paper introduces a method to develop phenotyping algorithms in an unbiased manner by automatically extracting and selecting informative features, which can be comparable to expert-curated ones in classification accuracy. MATERIALS AND METHODS: Comprehensive medical concepts were collected from publicly available knowledge sources in an automated, unbiased fashion. Natural language processing (NLP) revealed the occurrence patterns of these concepts in EHR narrative notes, which enabled selection of informative features for phenotype classification. When combined with additional codified features, a penalized logistic regression model was trained to classify the target phenotype. RESULTS: The authors applied our method to develop algorithms to identify patients with rheumatoid arthritis and coronary artery disease cases among those with rheumatoid arthritis from a large multi-institutional EHR. The area under the receiver operating characteristic curves (AUC) for classifying RA and CAD using models trained with automated features were 0.951 and 0.929, respectively, compared to the AUCs of 0.938 and 0.929 by models trained with expert-curated features. DISCUSSION: Models trained with NLP text features selected through an unbiased, automated procedure achieved comparable or slightly higher accuracy than those trained with expert-curated features. The majority of the selected model features were interpretable. CONCLUSION: The proposed automated feature extraction method, generating highly accurate phenotyping algorithms with improved efficiency, is a significant step toward high-throughput phenotyping.
Sheng Yu 0002, Katherine P. Liao, Stanley Y. Shaw, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai
J. Am. Medical Informatics Assoc.8
2014 Smart on FHIR
David McCallie, Joshua C. Mandel, Stanley M. Huff, Kenneth D. Mandl, Isaac S. Kohane
AMIA5
2014 Informatics for Integrating Biology and the Bedside (I2b2) Clinical Trials (CT) Patient Ascertainment Suite
Shawn N. Murphy, Nich Wattanasin, Susanne E. Churchill, Isaac S. Kohane, Vivian S. Gainer
AMIA4
2014 Improving the Review of Individual Patients in a Clinical Data Repository
Nich Wattanasin, Michael Mendis, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy
AMIA4
2014 Are Meaningful Use Stage 2 certified EHRs ready for interoperability? Findings from the SMART C-CDA Collaborative
abstract
BACKGROUND AND OBJECTIVE: Upgrades to electronic health record (EHR) systems scheduled to be introduced in the USA in 2014 will advance document interoperability between care providers. Specifically, the second stage of the federal incentive program for EHR adoption, known as Meaningful Use, requires use of the Consolidated Clinical Document Architecture (C-CDA) for document exchange. In an effort to examine and improve C-CDA based exchange, the SMART (Substitutable Medical Applications and Reusable Technology) C-CDA Collaborative brought together a group of certified EHR and other health information technology vendors. MATERIALS AND METHODS: We examined the machine-readable content of collected samples for semantic correctness and consistency. This included parsing with the open-source BlueButton.js tool, testing with a validator used in EHR certification, scoring with an automated open-source tool, and manual inspection. We also conducted group and individual review sessions with participating vendors to understand their interpretation of C-CDA specifications and requirements. RESULTS: We contacted 107 health information technology organizations and collected 91 C-CDA sample documents from 21 distinct technologies. Manual and automated document inspection led to 615 observations of errors and data expression variation across represented technologies. Based upon our analysis and vendor discussions, we identified 11 specific areas that represent relevant barriers to the interoperability of C-CDA documents. CONCLUSIONS: We identified errors and permissible heterogeneity in C-CDA documents that will limit semantic interoperability. Our findings also point to several practical opportunities to improve C-CDA document quality and exchange in the coming years.
John D. D'Amore, Joshua C. Mandel, David A. Kreda, Ashley Swain, George A. Koromia, Sumesh Sundareswaran, Liora Alschuler, Robert H. Dolin, Kenneth D. Mandl, Isaac S. Kohane, Rachel Badovinac Ramoni
J. Am. Medical Informatics Assoc.10
2014 Brief communication: Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS): Architecture
abstract
We describe the architecture of the Patient Centered Outcomes Research Institute (PCORI) funded Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS, http://www.SCILHS.org) clinical data research network, which leverages the $48 billion dollar federal investment in health information technology (IT) to enable a queryable semantic data model across 10 health systems covering more than 8 million patients, plugging universally into the point of care, generating evidence and discovery, and thereby enabling clinician and patient participation in research during the patient encounter. Central to the success of SCILHS is development of innovative 'apps' to improve PCOR research methods and capacitate point of care functions such as consent, enrollment, randomization, and outreach for patient-reported outcomes. SCILHS adapts and extends an existing national research network formed on an advanced IT infrastructure built with open source, free, modular components.
Kenneth D. Mandl, Isaac S. Kohane, Douglas MacFadden, Griffin M. Weber, Marc D. Natter, Joshua C. Mandel, Sebastian Schneeweiss, Sarah Weiler, Jeffrey G. Klann, Jonathan P. Bickel, William G. Adams, Yaorong Ge, James Perkins, Keith Marsolo, Elmer V. Bernstam, John Showalter, Alexander Quarshie, Elizabeth O. Ofili, George Hripcsak, Shawn N. Murphy
J. Am. Medical Informatics Assoc.2
2013 Organization and Transformation of Next-Generation Sequencing Data for Use within i2b2
Lori C. Phillips, Shawn N. Murphy, Isaac S. Kohane
AMIA3
2013 Integrating the CCDA for Real-Time Patient Data in the i2b2 Platform
Nich Wattanasin, Michael Mendis, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy
AMIA6
2012 Supporting Population Queries and Clinical Trials in i2b2 with SMART
Shawn N. Murphy, Michael Mendis, Nich Wattanasin, Alyssa Porter, Stella Ubaha, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane
AMIA10
2012 Apps to display patient data, making SMART available in the i2b2 platform
Nich Wattanasin, Alyssa Porter, Stella Ubaha, Michael Mendis, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy
AMIA9
2012 Integrating Substitutable Medical Apps, Reusable Technologies (SMART) in the i2b2 Platform
Nich Wattanasin, Alyssa Porter, Stella Ubaha, Michael Mendis, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy
AMIA9
2012 Quantifying the white blood cell transcriptome as an accessible window to the multiorgan transcriptome
abstract
Abstract Motivation: We investigate and quantify the generalizability of the white blood cell (WBC) transcriptome to the general, multiorgan transcriptome. We use data from the NCBI's Gene Expression Omnibus (GEO) public repository to define two datasets for comparison, WBC and OO (Other Organ) sets. Results: Comprehensive pair-wise correlation and expression level profiles are calculated for both datasets (with sizes of 81 and 1463, respectively). We have used mapping and ranking across the Gene Ontology (GO) categories to quantify similarity between the two sets. GO mappings of the most correlated and highly expressed genes from the two datasets tightly match, with the notable exceptions of components of the ribosome, cell adhesion and immune response. That is, 10 877 or 48.8% of all measured genes do not change >10% of rank range between WBC and OO; only 878 (3.9%) change rank >50%. Two trans-tissue gene lists are defined, the most changing and the least changing genes in expression rank. We also provide a general, quantitative measure of the probability of expression rank and correlation profile in the OO system given the expression rank and correlation profile in the WBC dataset. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Isaac S. Kohane, Vladimir I. Valtchinov
Bioinform.1
2012 Quantifying the white blood cell transcriptome as an accessible window to the multiorgan transcriptome
abstract
MOTIVATION: We investigate and quantify the generalizability of the white blood cell (WBC) transcriptome to the general, multiorgan transcriptome. We use data from the NCBI's Gene Expression Omnibus (GEO) public repository to define two datasets for comparison, WBC and OO (Other Organ) sets. RESULTS: Comprehensive pair-wise correlation and expression level profiles are calculated for both datasets (with sizes of 81 and 1463, respectively). We have used mapping and ranking across the Gene Ontology (GO) categories to quantify similarity between the two sets. GO mappings of the most correlated and highly expressed genes from the two datasets tightly match, with the notable exceptions of components of the ribosome, cell adhesion and immune response. That is, 10 877 or 48.8% of all measured genes do not change >10% of rank range between WBC and OO; only 878 (3.9%) change rank >50%. Two trans-tissue gene lists are defined, the most changing and the least changing genes in expression rank. We also provide a general, quantitative measure of the probability of expression rank and correlation profile in the OO system given the expression rank and correlation profile in the WBC dataset.
Isaac S. Kohane, Vladimir I. Valtchinov
Bioinform.1
2012 A translational engine at the national scale: informatics for integrating biology and the bedside
abstract
Informatics for integrating biology and the bedside (i2b2) seeks to provide the instrumentation for using the informational by-products of health care and the biological materials accumulated through the delivery of health care to conduct discovery research and to study the healthcare system in vivo. This complements existing efforts such as prospective cohort studies or trials outside the delivery of routine health care. i2b2 has been used to generate genome-wide studies at less than one tenth the cost and one tenth the time of conventionally performed studies as well as to identify important risk from commonly used medications. i2b2 has been adopted by over 60 academic health centers internationally.
Isaac S. Kohane, Susanne E. Churchill, Shawn N. Murphy
J. Am. Medical Informatics Assoc.1
2012 The SMART Platform: early experience enabling substitutable applications for electronic health records
abstract
OBJECTIVE: The Substitutable Medical Applications, Reusable Technologies (SMART) Platforms project seeks to develop a health information technology platform with substitutable applications (apps) constructed around core services. The authors believe this is a promising approach to driving down healthcare costs, supporting standards evolution, accommodating differences in care workflow, fostering competition in the market, and accelerating innovation. MATERIALS AND METHODS: The Office of the National Coordinator for Health Information Technology, through the Strategic Health IT Advanced Research Projects (SHARP) Program, funds the project. The SMART team has focused on enabling the property of substitutability through an app programming interface leveraging web standards, presenting predictable data payloads, and abstracting away many details of enterprise health information technology systems. Containers--health information technology systems, such as electronic health records (EHR), personally controlled health records, and health information exchanges that use the SMART app programming interface or a portion of it--marshal data sources and present data simply, reliably, and consistently to apps. RESULTS: The SMART team has completed the first phase of the project (a) defining an app programming interface, (b) developing containers, and (c) producing a set of charter apps that showcase the system capabilities. A focal point of this phase was the SMART Apps Challenge, publicized by the White House, using http://www.challenge.gov website, and generating 15 app submissions with diverse functionality. CONCLUSION: Key strategic decisions must be made about the most effective market for further disseminating SMART: existing market-leading EHR vendors, new entrants into the EHR market, or other stakeholders such as health information exchanges.
Kenneth D. Mandl, Joshua C. Mandel, Shawn N. Murphy, Elmer V. Bernstam, Rachel Badovinac Ramoni, David A. Kreda, J. Michael McCoy, Ben Adida, Isaac S. Kohane
J. Am. Medical Informatics Assoc.9
2012 Technical desiderata for the integration of genomic data into Electronic Health Records
Daniel R. Masys, Gail P. Jarvik, Neil F. Abernethy, Nicholas R. Anderson 0001, George J. Papanicolaou, Dina N. Paltoo, Mark A. Hoffman, Isaac S. Kohane, Howard P. Levy
J. Biomed. Informatics8
2011 Automated validation of genetic variants from large databases: ensuring that variant references refer to the same genomic locations
abstract
SUMMARY: Accurate annotations of genomic variants are necessary to achieve full-genome clinical interpretations that are scientifically sound and medically relevant. Many disease associations, especially those reported before the completion of the HGP, are limited in applicability because of potential inconsistencies with our current standards for genomic coordinates, nomenclature and gene structure. In an effort to validate and link variants from the medical genetics literature to an unambiguous reference for each variant, we developed a software pipeline and reviewed 68 641 single amino acid mutations from Online Mendelian Inheritance in Man (OMIM), Human Gene Mutation Database (HGMD) and dbSNP. The frequency of unresolved mutation annotations varied widely among the databases, ranging from 4 to 23%. A taxonomy of primary causes for unresolved mutations was produced. AVAILABILITY: This program is freely available from the web site (http://safegene.hms.harvard.edu/aa2nt/).
Mark Y. Tong, Christopher A. Cassa, Isaac S. Kohane
Bioinform.3
2011 BioNOT: A searchable database of biomedical negated sentences
abstract
BACKGROUND: Negated biomedical events are often ignored by text-mining applications; however, such events carry scientific significance. We report on the development of BioNØT, a database of negated sentences that can be used to extract such negated events. DESCRIPTION: Currently BioNØT incorporates ≈32 million negated sentences, extracted from over 336 million biomedical sentences from three resources: ≈2 million full-text biomedical articles in Elsevier and the PubMed Central, as well as ≈20 million abstracts in PubMed. We evaluated BioNØT on three important genetic disorders: autism, Alzheimer's disease and Parkinson's disease, and found that BioNØT is able to capture negated events that may be ignored by experts. CONCLUSIONS: The BioNØT database can be a useful resource for biomedical researchers. BioNØT is freely available at http://bionot.askhermes.org/. In future work, we will develop semantic web related technologies to enrich BioNØT.
Shashank Agarwal, Hong Yu 0001, Isaac S. Kohane
BMC Bioinform.3
2011 Marco Ramoni: an appreciation of academic achievement
abstract
We review the scholarly career of our colleague, Marco Ramoni, who died unexpectedly in the summer of 2010. His work mainly explored the development and application of Bayesian techniques to model clinical, public health, and bioinformatics questions. His contributions have led to improvements in our ability to model behavior that evolves in time, to explore systematic relationships among large sets of covariates, and to tease out the meaning of data on the role of genetic variation in the genesis of important diseases.
Isaac S. Kohane, Peter Szolovits
J. Am. Medical Informatics Assoc.1
2011 Strategies for maintaining patient privacy in i2b2
abstract
BACKGROUND: The re-use of patient data from electronic healthcare record systems can provide tremendous benefits for clinical research, but measures to protect patient privacy while utilizing these records have many challenges. Some of these challenges arise from a misperception that the problem should be solved technically when actually the problem needs a holistic solution. OBJECTIVE: The authors' experience with informatics for integrating biology and the bedside (i2b2) use cases indicates that the privacy of the patient should be considered on three fronts: technical de-identification of the data, trust in the researcher and the research, and the security of the underlying technical platforms. METHODS: The security structure of i2b2 is implemented based on consideration of all three fronts. It has been supported with several use cases across the USA, resulting in five privacy categories of users that serve to protect the data while supporting the use cases. RESULTS: The i2b2 architecture is designed to provide consistency and faithfully implement these user privacy categories. These privacy categories help reflect the policy of both the Health Insurance Portability and Accountability Act and the provisions of the National Research Act of 1974, as embodied by current institutional review boards. CONCLUSION: By implementing a holistic approach to patient privacy solutions, i2b2 is able to help close the gap between principle and practice.
Shawn N. Murphy, Vivian S. Gainer, Michael Mendis, Susanne E. Churchill, Isaac S. Kohane
J. Am. Medical Informatics Assoc.5
2011 Inferring cell cycle feedback regulation from gene expression data
Fulvia Ferrazzi, Felix B. Engel, Erxi Wu, Annie P. Moseman, Isaac S. Kohane, Riccardo Bellazzi, Marco Ramoni
J. Biomed. Informatics5
2010 Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2)
abstract
Informatics for Integrating Biology and the Bedside (i2b2) is one of seven projects sponsored by the NIH Roadmap National Centers for Biomedical Computing (http://www.ncbcs.org). Its mission is to provide clinical investigators with the tools necessary to integrate medical record and clinical research data in the genomics age, a software suite to construct and integrate the modern clinical research chart. i2b2 software may be used by an enterprise's research community to find sets of interesting patients from electronic patient medical record data, while preserving patient privacy through a query tool interface. Project-specific mini-databases ("data marts") can be created from these sets to make highly detailed data available on these specific patients to the investigators on the i2b2 platform, as reviewed and restricted by the Institutional Review Board. The current version of this software has been released into the public domain and is available at the URL: http://www.i2b2.org/software.
Shawn N. Murphy, Griffin M. Weber, Michael Mendis, Vivian S. Gainer, Henry C. Chueh, Susanne E. Churchill, Isaac S. Kohane
J. Am. Medical Informatics Assoc.7
2009 Integration of heterogeneous expression data sets extends the role of the retinol pathway in diabetes and insulin resistance
abstract
MOTIVATION: Type 2 diabetes is a chronic metabolic disease that involves both environmental and genetic factors. To understand the genetics of type 2 diabetes and insulin resistance, the DIabetes Genome Anatomy Project (DGAP) was launched to profile gene expression in a variety of related animal models and human subjects. We asked whether these heterogeneous models can be integrated to provide consistent and robust biological insights into the biology of insulin resistance. RESULTS: We perform integrative analysis of the 16 DGAP data sets that span multiple tissues, conditions, array types, laboratories, species, genetic backgrounds and study designs. For each data set, we identify differentially expressed genes compared with control. Then, for the combined data, we rank genes according to the frequency with which they were found to be statistically significant across data sets. This analysis reveals RetSat as a widely shared component of mechanisms involved in insulin resistance and sensitivity and adds to the growing importance of the retinol pathway in diabetes, adipogenesis and insulin resistance. Top candidates obtained from our analysis have been confirmed in recent laboratory studies.
Peter J. Park, Sek Won Kong, Toma Tebaldi, Weil R. Lai, Simon Kasif, Isaac S. Kohane
Bioinform.6
2009 Research Paper: Prediction of Chronic Obstructive Pulmonary Disease (COPD) in Asthma Patients Using Electronic Medical Records
abstract
OBJECTIVE: Identify clinical factors that modulate the risk of progression to COPD among asthma patients using data extracted from electronic medical records. DESIGN: Demographic information and comorbidities from adult asthma patients who were observed for at least 5 years with initial observation dates between 1988 and 1998, were extracted from electronic medical records of the Partners Healthcare System using tools of the National Center for Biomedical Computing "Informatics for Integrating Biology to the Bedside" (i2b2). MEASUREMENTS: A predictive model of COPD was constructed from a set of 9,349 patients (843 cases, 8,506 controls) using Bayesian networks. The model's predictive accuracy was tested using it to predict COPD in a future independent set of asthma patients (992 patients; 46 cases, 946 controls), who had initial observation dates between 1999 and 2002. RESULTS: A Bayesian network model composed of age, sex, race, smoking history, and 8 comorbidity variables is able to predict COPD in the independent set of patients with an accuracy of 83.3%, computed as the area under the Receiver Operating Characteristic curve (AUROC). CONCLUSIONS: Our results demonstrate that data extracted from electronic medical records can be used to create predictive models. With improvements in data extraction and inclusion of more variables, such models may prove to be clinically useful.
Blanca E. Himes, Isaac S. Kohane, Scott T. Weiss, Marco Ramoni
J. Am. Medical Informatics Assoc.3
2009 Application of Information Technology: The Shared Health Research Information Network (SHRINE): A Prototype Federated Query Tool for Clinical Data Repositories
abstract
The authors developed a prototype Shared Health Research Information Network (SHRINE) to identify the technical, regulatory, and political challenges of creating a federated query tool for clinical data repositories. Separate Institutional Review Boards (IRBs) at Harvard's three largest affiliated health centers approved use of their data, and the Harvard Medical School IRB approved building a Query Aggregator Interface that can simultaneously send queries to each hospital and display aggregate counts of the number of matching patients. Our experience creating three local repositories using the open source Informatics for Integrating Biology and the Bedside (i2b2) platform can be used as a road map for other institutions. The authors are actively working with the IRBs and regulatory groups to develop procedures that will ultimately allow investigators to obtain identified patient data and biomaterials through SHRINE. This will guide us in creating a future technical architecture that is scalable to a national level, compliant with ethical guidelines, and protective of the interests of the participating hospitals.
Griffin M. Weber, Shawn N. Murphy, Andrew J. McMurry, Douglas MacFadden, Daniel J. Nigrin, Susanne E. Churchill, Isaac S. Kohane
J. Am. Medical Informatics Assoc.7
2008 Characterization of Patients who Suffer Asthma Exacerbations using Data Extracted from Electronic Medical Records
Blanca E. Himes, Isaac S. Kohane, Marco Ramoni, Scott T. Weiss
AMIA2
2008 Letters to the Editor: No Structure Before Its Time
abstract
Dr. Schleyer asks a number of important questions. These might be summarized as asking why, as a discipline, are we not focusing on improving the acquisition of structured data rather than going through computational acrobatics to extract codified representation from narrative text? Should we not be focusing our efforts to ensure a fully-structured record? We agree with Dr. Schleyer that a fully-structured record is an important goal. In the context of a healthcare system with 10-minute visits, time of data entry using current technologies remains burdensome.1 Even if time of data entry were of no concern, the engineering of structured data acquisition interfaces to support the full richness of patient state captured by natural language remains a challenge. Most tellingly, there have been over four decades of research by leading informaticians in precisely the area that Dr. Schleyer urges us to explore and yet the volume of narrative text continues to grow in our leading academic centers. Perhaps the greatest potential for developing a fully-structured medical record lies in patients' annotating their own record given their direct interest in accuracy and disposable time. Forty years ago, it was shown that patients could accurately present their symptoms in codified fashion2 and subsequent research suggests that they and their families can do equally well in reporting medications and even physiological state.3,4 With the increased popularity of personally-controlled health records, such capabilities are likely to be increasingly exploited. Notwithstanding these hopeful developments, it seems quite likely that the amount of narrative text describing patient states will continue to grow in the near to mid-term and the existing corpus will persist for at least the life of our patients. The obligation and the opportunity to provide the best possible care of our patients and the most informed research using such data will therefore continue to motivate segments of the academic and commercial informatics research communities in refining natural language processing techniques.
Isaac S. Kohane, Özlem Uzuner
J. Am. Medical Informatics Assoc.1
2008 Viewpoint Paper: Identifying Patient Smoking Status from Medical Discharge Records
abstract
Clinical narrative records contain much useful information. However, most clinical narratives are in the form of fragmented English free text, showing the characteristics of a clinical sublanguage. This makes their linguistic processing, search, and retrieval challenging.1 Traditional natural language processing (NLP) tools are not designed for the fragmented free text found in narrative clinical records; therefore, they do not perform well on this type of data.2 Limited access to clinical records has been a barrier to the widespread development of medical language processing (MLP) technologies. In the absence of a standardized, publicly available ground truth that encourages the development of MLP systems and allows their head-to-head comparison, successful MLP efforts have been limited, e.g., MedLEE3 and Symtxt.4 A few MLP systems have been developed,5 and such efforts have successfully shown the usefulness of MLP in clinical settings.6–8 To improve the availability of clinical records and to contribute to the advancement of the state of the art in MLP, within the i2b2 (Informatics for Integrating Biology to the Bedside) project, the authors de-identified and released a set of clinical records from Partners HealthCare. These records provided the basis for the development of ground truth for two challenge questions: Automatic de-identification of clinical data, i.e., de-identification challenge. Automatic evaluation of the smoking status of patients based on medical records, i.e., smoking challenge. Representative teams from the MLP community participated in the two challenges and met at a workshop organized by the authors to discuss the results of the challenges. The workshop was co-sponsored by the American Medical Informatics Association and met in conjunction with its Fall Symposium in November 2006. This article provides an overview of the smoking challenge and the findings of the workshop. An overview of the de-identification challenge can be found in Uzuner et al.9 The smoking challenge continues the tradition of attempting to identify the state of the art in automatic language processing. Outside of the medical domain, there have been many efforts in this direction. Most of these efforts have been led by Message Understanding Conferences (MUC)10 and the National Institute of Standards and Technology (NIST).11 MUC organized shared tasks on named entity recognition. NIST organized a series of Text Retrieval Evaluation Conferences (TREC) on various domains including blogs and legal documents; they also organized a series of shared tasks on topic detection and tracking, speaker recognition, language recognition, spoken document retrieval, machine translation, and entity extraction. In the biomedical domain, three such prominent efforts were BioCreAtIvE12 for information extraction, ImageCLEF13,14 for image retrieval, and TREC Genomics15 for question answering and information retrieval. Keeping the goals of TREC,16 MUC,17 BioCreAtIvE,18 etc., in mind for the smoking challenge, we created a collection of actual medical discharge records. We invited the development of systems that can predict the smoking status of patients based on the narratives in these medical discharge records. We limited the scope of this task to understanding only the explicitly reported smoking information. In other words, information that implicitly reveals the smoking status was excluded from this study. Our smoking challenge continued the work on the application of classification techniques to the medical domain,19–22 and extended the MLP studies on medical discharge records.8,23–29 Information on the smoking status of patients is important for many health studies, e.g., studies on asthma; however, before this challenge, the only system for the automatic evaluation of the smoking status of patients from their records was the HITEx system.30 The data for the smoking challenge consisted exclusively of discharge summaries from Partners HealthCare. We preprocessed these records so that they were de-identified, tokenized, broken into sentences, converted into XML format, and separated into training and test sets. Institutional review boards of Partners HealthCare, Massachusetts Institute of Technology, and the State University of New York at Albany approved the challenge and the data preparation process. The data for the challenge were annotated by pulmonologists. The pulmonologists were asked to classify patient records into five possible smoking status categories. For the purposes of this challenge, we defined these categories as follows: A Past Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker one year or more ago but who has not smoked for at least one year. The assertion “past smoker” without any temporal qualifications means Past Smoker unless there is text that says that the patient stopped smoking less than one year ago. A Current Smoker is a patient whose discharge summary asserts explicitly that the patient was a smoker within the past year. The assertion “current smoker” without any temporal qualifications means Current Smoker unless there is text that says that the patient stopped smoking more than a year ago. A Smoker is a patient who is either a Current or a Past Smoker but whose medical record does not provide enough information to classify the patient as either. A Non-Smoker's discharge summary indicates that they never smoked. An Unknown is a patient whose discharge summary does not mention anything about smoking. Indecision between Current Smoker and Past Smoker does not belong to this category. Second-hand smokers are considered Non-Smokers for the purposes of this study, unless there is evidence in their record that they actively smoked. Similarly, as we are only concerned with tobacco, marijuana smoking should not affect the patients' smoking status. In addition to being provided with the above definitions, the annotators were trained on 55 sample sentences30 (see Table 1 for a subset) and 10 sample records. Annotator Training Samples Annotator Training Samples Two pulmonologists annotated each record with the smoking status of patients based strictly on the explicitly stated smoking-related facts in the records. These annotations constitute the textual judgments of the annotators. The same two pulmonologists also marked the smoking status of the patients using their medical intuitions on all information in the records. These annotations constitute the intuitive judgments of the annotators. In all, 928 records were annotated. The interannotator agreement on the textual judgments on these records, as measured by Cohen's kappa (κ),31,32 was 0.84; observed agreement on textual judgments was 0.93; specific agreement per category on textual judgments ranged from 0.4 to 0.98 (Table 2; also see the Methods section for definitions of Cohen's kappa, observed agreement, and specific agreement). The interannotator agreement on the intuitive judgments was 0.45; observed agreement on intuitive judgments was 0.73; specific agreement per category on intuitive judgments ranged from 0.3 to 0.84 (Table 2). Observed and Specific Agreement Observed and Specific Agreement Guidelines for κ are subject to interpretation and depend on parameters such as the task and categories involved.33 However, κ of 0.8 is widely used as the threshold for strong agreement.31–35 On our data, we observed strong agreement only on the textual judgments. We further observed that the intra-annotator agreement between a given doctor's intuitive and textual judgments varied from 0.62 to 0.99. This indicates that the reliance of the intuitive judgments on the explicit textual information varies considerably from doctor to doctor. Given these observations, we limited the challenge task to the identification of the smoking status based on information that is explicitly mentioned in the records. To generate the ground truth, we resolved the disagreements (on textual judgments) between the annotators by obtaining judgments from two other pulmonologists. We omitted the records that the annotators disagreed on from the challenge, unless a majority vote could identify a clear textual judgment for them. In all, 63 records were omitted from the challenge for lack of a clear textual judgment. In addition, annotation results showed heavy bias in the data for Unknown records. This is the least interesting category for the purposes of the smoking challenge as Unknown records do not contain any smoking-related information. To focus the smoking challenge less on the Unknown category and more on the other four categories, we omitted a portion of the Unknown records (363 records) from the challenge. A total of 502 de-identified medical discharge records were used for the smoking challenge. Table 3 shows the distribution of annotated records into training and test sets, and into Past Smoker, Current Smoker, Smoker, Non-Smoker, and Unknown categories. The training and test sets show similar distribution of records into the five smoking categories; however, these distributions are far from uniform. This reflects the realities of real-world data; our records were drawn at random from the Partners' database, in which some smoking categories are better represented than others. In our test set, the smallest smoking category is Smokers, with only three records. The training and test data can be obtained from i2b2.org. Smoking Status Training and Test Data Distribution Smoking Status Training and Test Data Distribution We evaluated system performances using microaveraged and macroaveraged precision, recall, and F-measure, as well as Cohen's kappa. Precision, recall, and F-measure are performance metrics frequently used in NLP.36,37 These metrics are easily derived from a binary confusion matrix. In a binary decision problem, a classifier labels entities as either positive or negative (where positive and negative represent two generic categories) and produces a confusion matrix. This matrix contains four entities: true positive (TP), true negative (TN), false positive (FP), and false negative (FN). Given such a matrix, precision is the percentage of entities classified correctly to be in a given category in relation to the total number of entities classified for the given category (Equation 1). Recall is the percentage of entities classified correctly in a given category in relation to the actual number of items in the given category (Equation 2). F-measure is the harmonic mean of precision and recall (Equation 3). β enables F-measure to favor either precision or recall. We give equal weight to precision and recall by setting β = 1. Cohen's kappa (κ) (Equation 6) is a measure of agreement31 between pairs of annotators who classify items into a set number of mutually exclusive categories. κ depends on observed agreement (Ao in Equation 7) and the agreement expected due to chance (Ae in Equation 8). A κ value of 0.8 is widely used as the threshold for strong agreement,31–35 whereas a κ of 0 indicates that the observed agreement is due to chance.33 We used κ as a measure of inter-annotator and intra-annotator agreement (see Annotations section) as a measure of agreement between two automatic systems and as a measure of agreement between a system and the ground truth (see Results and Discussion). Equation 6 through Equation 8 collectively describe κ between an automatic system and the ground truth. κ for inter-annotator and intra-annotator agreement and for inter-system agreement can be computed analogously. According to Hripcsak and Rothschild,38 there exists a correspondence between F-measure and κ. However, κ provides clearer insights into the relative strengths of the systems (see Intersystem Agreement section). We evaluated systems using F-measure but compared them using both F-measure and κ. Specific agreement (Asp)33 measures the degree of agreement (Equation 9) on each category and is not adjusted by chance. We used specific agreement to get a sense of the level of agreement between annotators without taking chance into consideration. We tested the significance of the differences of the systems using a randomization technique that is frequently utilized in NLP.39 The null hypothesis is that the absolute value of the difference in performances, e.g., F-measures, of two systems is approximately equal to zero. The randomization technique does not assume a particular distribution of the differences. Instead, it empirically generates the distribution. Given two actual systems, it randomly shuffles (at each iteration, we simulated a coin flip to decide whether the answers should be swapped) their responses to the records in the test set N times (e.g., N = 9,999), and thus creates N pairs of pseudosystems. It counts the number of times that the difference between the performances of pairs of pseudosystems is greater than the difference between the two actual systems' performances. Let this count be equal to n and compute . If s is greater than a predetermined cutoff α, then the difference of the performances of the two actual systems can be explained by chance; otherwise, the difference is significant at level α. Following MUC's example, we set α to 0.1. A total of 11 teams participated in the smoking challenge. The training data for the challenge were released in July 2006, and the test data were released for only three days in September 2006. Each team was permitted to submit up to three system on the test A total of were count only one of the three of and their other two were evaluated from the (see and are not in this In this we describe each et a this classifier that to the smoking status of patients and then and to the information. The of the that only a few in a record to the smoking status and that these could be easily by their (e.g., If more than one in a record smoking then only the was If such were then the record was classified as To predict the smoking status of a record from the test set, each from this record was compared with from the training The of the measures between each and the most similar in the training set the smoking status of the et various for text with and found that the performance from the of and that and classification more than et also from lack of explicit smoking information in the Unknown and these to further two to processing the In the they classified based on to smoking using In the they classified each explicit to smoking using and then to categories from judgments of to smoking the smoking category for the collection the category in the Given the data of their et the i2b2 data set with records and their smoking categories. that a based on the data set perform better than trained on each of the data sets This hypothesis is with a the of a system is as the sample et smoking status evaluation system was with and from the Medical This system document (e.g., medical and the status of these entities (e.g., smoking-related such as and It thus the set of and of explicit to smoking. et found that the using on the to smoking better than the classification using (see and in Table and Table this system are by the of the and for example, to the system by et The and of explicit to their For example, the does not in the section could both to a Past Smoker and to a Non-Smoker, on However, such and information was to and for Precision, and by and for Precision, and by on and by not in macroaveraged not in microaveraged the is on and by not in macroaveraged not in microaveraged the is smoking status evaluation as a task and a process. The marked smoking-related In Cohen's a smoking-related was a of specific (e.g., and The the with the from the The without any specific and them The the and classified them using created of system by a system of for the smoking status of This system utilized both and such as of and of three one and two for the smoking status of patients of smoking challenge systems as a to the and can be at For with from the and found decision the most evaluated using on the training this classifier its performance trained using only the and their as For with and were with up to five words, one of the two in the with the These two systems to most of the test data to the Unknown category. found that of as well as Table shows that from in both microaveraged and macroaveraged F-measures, at α = 0.1. the Medical to the challenge. is a system for medical text processing in and records by For the smoking challenge, this system was to English medical discharge et the smoking challenge as a classification system marked the smoking status category of each in a record and to classify the document The judgments of the training set were derived Most of the text processing was through the of the Information system of and classification was in three and by using with an In the the Unknown category was a of that excluded this category. In the the category was through the of For with was In the the Smoker, and Past Smoker categories were by a of to Current and Past This was as a temporal and was using an the in this the were not so as to information of such as the of the and such as and between Current and Past the categories were The document categories were from to as follows: Current Smoker, Past Smoker, Smoker, Non-Smoker, and Each document was the of the category for which it could provide For example, a Current Smoker was to a document of any as Current et and the system to the smoking challenge. is designed for and clinical information from free text, and contains a patient smoking status This system of four document and The and the of a record based on the of the section The to them with in a The the text and is to The to labels to each The set of is then using provides based on of patient This system for a given patient and a to the with the to each To the differences of discharge summaries from patient et to the smoking challenge. this they into the differences in the labels of the i2b2 smoking challenge and the smoking they the system to a specific than a set of and for each The reported that they found the temporal of smoking categories to be a significant challenge. et of the explicit to smoking status. and using the results of these and records using a In addition to they used and information about and used a system to predict the smoking status based on explicit of smoking. that the system they for information about the smoking status of patients the records were of explicit smoking-related For they all explicit of smoking from the records and created a data set that only the trained two on the showed that annotators precision, recall, and F-measure on the The two systems the performance of annotators with of on the test Table and 1 show the precision, recall, and F-measure, both macroaveraged and for each of the system to the smoking challenge. Table shows the results of the significance on microaveraged and macroaveraged of the this reveals that the differences in the microaveraged of the systems are not significant at α = of these systems are not from each other in their macroaveraged at this α. Results from Table by microaveraged Table shows that systems that used similar machine e.g., the systems that used not all perform This that other such as the and the used with the also to the Table also shows that a majority of the performances from systems that of the lack of explicit smoking information in the Unknown records and to further processing or classification (see the results of et and et These systems microaveraged to that were on explicit to smoking status and et of the systems that the smoking classification was Unknown unless some in the document showed it to be this in an F-measure of in the Unknown category for and et and et Table 6 shows the precision, recall, and F-measure of each system on each of the five smoking status categories. In the systems successfully the Unknown and categories; some Current and Past with the of the systems of et systems correctly classified one of the three all in Precision, and for in Precision, and for in In addition to system performance on the test set and on categories in the test set, we the performance and agreement of systems on data i.e., records, in the test For we used κ. Table shows the level of κ agreement of each system with the ground truth as well as the level of κ agreement of pairs of Most is the level of agreement between systems by the same but only for some of the This level of agreement is not these from of the same The systems of et in their to and and only in their training and used a that the from and Agreement of and the Agreement between 0.8 and is in and agreement is in Agreement of and the Agreement between 0.8 and is in and agreement is in Table also shows in the systems showed a level of agreement with each For example, each of the systems of et with each of Cohen's systems showed κ between and systems showed strong agreement with each other the differences in their to smoking status For example, and system showed κ of to compared with et Similarly, the systems of et a κ of 0.8 with that of and the classified the records at the level whereas the was at the The level of agreement between to the smoking challenge, by the of these on the ground truth, indicates the of the set of to smoking status On the other the disagreements the systems their relative For example, and disagreed with each other on records = These systems showed relative strengths in in and Current Smoker, and in Unknown and Past were only three records that system marked The state of the art in smoking status evaluation be by the strengths of such of the teams with the annotation of a few of the in the challenge mentioned in the Annotations our medical records were annotated by pulmonologists. they provide a ground truth, judgments can However, given the medical of the annotators and the agreement on the labels of these records the we the annotations as the English of the records is for all of the a majority of the systems found the classification disagreed with the judgments of we that of should be by an to measure the of the that the in the of the ground truth only on our findings on the smoking challenge, we to our in two we are by the strengths of the systems and to of these to improve the we and with data to be We that a system provide insights into the intuitive judgments on the smoking status of In this we the i2b2 smoking challenge, the data and the data preparation the evaluation and each of the We the evaluation and of the system and provided a of the of our findings for For this challenge, we a document collection derived from actual medical discharge records, this collection in to a real-world medical classification and system performances. We showed that asked to a decision on the smoking status of patients based on the explicitly stated information in medical discharge annotators with each other more than of the The systems that participated in the smoking challenge represented various from and to the differences in their to smoking status many of these systems In there were system with microaveraged above A majority of these systems of the of the challenge data, e.g., lack of to smoking in records marked of explicit to smoking. the systems in the smoking challenge showed that discharge summaries smoking status using a limited number of textual (e.g., of the smoking status from these The authors all teams for their to the challenge, for their in the of the workshop that the challenge, and and the of for their on this
Özlem Uzuner, Ira Goldstein, Yuan Luo 0001, Isaac S. Kohane
J. Am. Medical Informatics Assoc.4
2007 Architecture of the Open-source Clinical Research Chart from Informatics for Integrating Biology and the Bedside
Shawn N. Murphy, Michael Mendis, Kristel Hackett, Rajesh Kuttan, Wensong Pan, Lori C. Phillips, Vivian S. Gainer, David Berkowicz, John P. Glaser, Isaac S. Kohane, Henry C. Chueh
AMIA10
2007 Model Formulation: A Self-scaling, Distributed Information Architecture for Public Health, Research, and Clinical Care
abstract
OBJECTIVE: This study sought to define a scalable architecture to support the National Health Information Network (NHIN). This architecture must concurrently support a wide range of public health, research, and clinical care activities. STUDY DESIGN: The architecture fulfils five desiderata: (1) adopt a distributed approach to data storage to protect privacy, (2) enable strong institutional autonomy to engender participation, (3) provide oversight and transparency to ensure patient trust, (4) allow variable levels of access according to investigator needs and institutional policies, (5) define a self-scaling architecture that encourages voluntary regional collaborations that coalesce to form a nationwide network. RESULTS: Our model has been validated by a large-scale, multi-institution study involving seven medical centers for cancer research. It is the basis of one of four open architectures developed under funding from the Office of the National Coordinator of Health Information Technology, fulfilling the biosurveillance use case defined by the American Health Information Community. The model supports broad applicability for regional and national clinical information exchanges. CONCLUSIONS: This model shows the feasibility of an architecture wherein the requirements of care providers, investigators, and public health authorities are served by a distributed model that grants autonomy, protects privacy, and promotes participation.
Andrew J. McMurry, Clint A. Gilbert, Ben Y. Reis, Henry C. Chueh, Isaac S. Kohane, Kenneth D. Mandl
J. Am. Medical Informatics Assoc.5
2006 Mapping Tool for Maintaining Vocabulary Relationships
Kristel Hackett, Isaac S. Kohane, Henry C. Chueh, Shawn N. Murphy
AMIA2
2006 Integration of Clinical and Genetic Data in the i2b2 Architecture
Shawn N. Murphy, Michael Mendis, David A. Berkowitz, Isaac S. Kohane, Henry C. Chueh
AMIA4
2006 Integration of the Personally Controlled Electronic Medical Record into a Regional and Data Exchange: A National Demonstration
William W. Simons, John D. Halamka, Isaac S. Kohane, Daniel J. Nigrin, Nathan Finstein, Kenneth D. Mandl
AMIA3
2006 Dynamic Bayesian Networks in Modelling Cellular Systems: a Critical Appraisal on Simulated Data
abstract
Dynamic Bayesian networks offer a powerful modelling tool to unravel cellular mechanisms. In particular, Gaussian networks have recently been used to model gene expression data, thanks to their capability to avoid information loss associated with discretization and their good computational efficiency. Gaussian networks typically describe the conditional mean of a node as a linear regression of the parent variables. Such model can be generalized by using a linear regression of nonlinear transformations of the parent values. In this paper we investigate the use of both models and evaluate the performance of Gaussian networks in learning the complex dynamic interactions among genes and proteins. To this aim, we analyzed simulated data produced by a mathematical model of cell cycle control in budding yeast. The results obtained allowed us to appraise the performance of the different models and confirmed the suitability of dynamic Bayesian networks for a first level, genome-wide analysis of high throughput dynamic data
Fulvia Ferrazzi, Paola Sebastiani, Isaac S. Kohane, Marco Ramoni, Riccardo Bellazzi
CBMS3
2006 START: an automated tool for serial analysis of chromatin occupancy data
abstract
UNLABELLED: The serial analysis of chromatin occupancy technique (SACO) promises to become a widely used method for the unbiased genome-wide experimental identification of loci bound by a transcription factor of interest. We describe the first web-based automatic tool, termed sequence tag analysis and reporting tool (START), for processing SACO data generated by experiments performed for the yeast, fruit fly, mouse, rat or human genomes. The program uses as input sequences of inserts from a SACO library from which it extracts all SACO tags, maps them to genomic locations and annotates them. START returns detailed information about these tags including the genes, the genomic elements and the miRNA precursors found in their vicinity, and makes use of the MAPPER database to identify putative transcription factor binding sites located close to the tags. AVAILABILITY: The program is available at http://bio.chip.org/start/. SUPPLEMENTARY INFORMATION: SUPPLEMENTARY INFORMATION is available at http://bio.chip.org/doc/start/START-supplementary.pdf
Voichita D. Marinescu, Isaac S. Kohane, David A. Harmin, Michael E. Greenberg, Alberto Riva
Bioinform.2
2006 Lower expression of genes near microRNA in C. elegans germline
abstract
BACKGROUND: MicroRNAs (miRNAs) are recently discovered short non-protein-coding RNA molecules. miRNAs are increasingly implicated in tissue-specific transcriptional control and particularly in development. Because there is mounting evidence for the localized component of transcriptional control, we investigated if there is a distance-dependent effect of miRNA. RESULTS: We analyzed gene expression levels around the 84 of 113 know miRNAs for which there are nearby gene that were measured in the data in two independent C. elegans expression data sets. The expression levels are lower for genes in the vicinity of 59 of 84 (71%) miRNAs as compared to genes far from such miRNAs. Analysis of the genes with lower expression in proximity to the miRNAs reveals increased frequency matching of the 7 nucleotide "seed"s of these miRNAs. CONCLUSION: We found decreased messenger RNA (mRNA) abundance, localized within a 10 kb of chromosomal distance of some miRNAs, in C. elegans germline. The increased frequency of seed matching near miRNA can explain, in part, the localized effects.
Hidenori Inaoka, Yutaka Fukuoka, Isaac S. Kohane
BMC Bioinform.3
2005 Reverse Geocoding: Concerns about Patient Confidentiality in the Display of Geospatial Health Data
John S. Brownstein, Christopher A. Cassa, Isaac S. Kohane, Kenneth D. Mandl
AMIA3
2005 DITTO - a Tool for Identification of Patient Cohorts from the Text of Physician Notes in the Electronic Medical Record
Alexander Turchin, Merri L. Pendergrass, Isaac S. Kohane
AMIA3
2005 CrossChip: a system supporting comparative analysis of different generations of Affymetrix arrays
abstract
SUMMARY: To increase compatibility between different generations of Affymetrix GeneChip arrays, we propose a method of filtering probes based on their sequences. Our method is implemented as a web-based service for downloading necessary materials for converting the raw data files (*.CEL) for comparative analysis. The user can specify the appropriate level of filtering by setting the criteria for the minimum overlap length between probe sequences and the minimum number of usable probe pairs per probe set. Our website supports a within-species comparison for human and mouse GeneChip arrays. AVAILABILITY: http://www.crosschip.org
Sek Won Kong, Kyu-Baek Hwang, Richard D. Kim, Byoung-Tak Zhang, Steven A. Greenberg, Isaac S. Kohane, Peter J. Park
Bioinform.6
2005 Redefinition of Affymetrix probe sets by sequence overlap with cDNA microarray probes reduces cross-platform inconsistencies in cancer-associated gene expression measurements
abstract
BACKGROUND: Comparison of data produced on different microarray platforms often shows surprising discordance. It is not clear whether this discrepancy is caused by noisy data or by improper probe matching between platforms. We investigated whether the significant level of inconsistency between results produced by alternative gene expression microarray platforms could be reduced by stringent sequence matching of microarray probes. We mapped the short oligo probes of the Affymetrix platform onto cDNA clones of the Stanford microarray platform. Affymetrix probes were reassigned to redefined probe sets if they mapped to the same cDNA clone sequence, regardless of the original manufacturer-defined grouping. The NCI-60 gene expression profiles produced by Affymetrix HuFL platform were recalculated using these redefined probe sets and compared to previously published cDNA measurements of the same panel of RNA samples. RESULTS: The redefined probe sets displayed a substantially higher level of cross-platform consistency at the level of gene correlation, cell line correlation and unsupervised hierarchical clustering. The same strategy allowed an almost complete correspondence of breast cancer subtype classification between Affymetrix gene chip and cDNA microarray derived gene expression data, and gave an increased level of similarity between normal lung derived gene expression profiles using the two technologies. In total, two Affymetrix gene-chip platforms were remapped to three cDNA platforms in the various cross-platform analyses, resulting in improved concordance in each case. CONCLUSION: We have shown that probes which target overlapping transcript sequence regions on cDNA microarrays and Affymetrix gene-chips exhibit a greater level of concordance than the corresponding Unigene or sequence matched features. This method will be useful for the integrated analysis of gene expression data generated by multiple disparate measurement platforms.
Scott L. Carter, Aron C. Eklund, Brigham H. Mecham, Isaac S. Kohane, Zoltan Szallasi
BMC Bioinform.4
2005 MAPPER: a search engine for the computational identification of putative transcription factor binding sites in multiple genomes
abstract
BACKGROUND: Cis-regulatory modules are combinations of regulatory elements occurring in close proximity to each other that control the spatial and temporal expression of genes. The ability to identify them in a genome-wide manner depends on the availability of accurate models and of search methods able to detect putative regulatory elements with enhanced sensitivity and specificity. RESULTS: We describe the implementation of a search method for putative transcription factor binding sites (TFBSs) based on hidden Markov models built from alignments of known sites. We built 1,079 models of TFBSs using experimentally determined sequence alignments of sites provided by the TRANSFAC and JASPAR databases and used them to scan sequences of the human, mouse, fly, worm and yeast genomes. In several cases tested the method identified correctly experimentally characterized sites, with better specificity and sensitivity than other similar computational methods. Moreover, a large-scale comparison using synthetic data showed that in the majority of cases our method performed significantly better than a nucleotide weight matrix-based method. CONCLUSION: The search engine, available at http://mapper.chip.org, allows the identification, visualization and selection of putative TFBSs occurring in the promoter or other regions of a gene from the human, mouse, fly, worm and yeast genomes. In addition it allows the user to upload a sequence to query and to build a model by supplying a multiple sequence alignment of binding sites for a transcription factor of interest. Due to its extensive database of models, powerful search engine and flexible interface, MAPPER represents an effective resource for the large-scale computational analysis of transcriptional regulation.
Voichita D. Marinescu, Isaac S. Kohane, Alberto Riva
BMC Bioinform.2
2005 Systematic survey reveals general applicability of "guilt-by-association" within gene coexpression networks
abstract
BACKGROUND: Biological processes are carried out by coordinated modules of interacting molecules. As clustering methods demonstrate that genes with similar expression display increased likelihood of being associated with a common functional module, networks of coexpressed genes provide one framework for assigning gene function. This has informed the guilt-by-association (GBA) heuristic, widely invoked in functional genomics. Yet although the idea of GBA is accepted, the breadth of GBA applicability is uncertain. RESULTS: We developed methods to systematically explore the breadth of GBA across a large and varied corpus of expression data to answer the following question: To what extent is the GBA heuristic broadly applicable to the transcriptome and conversely how broadly is GBA captured by a priori knowledge represented in the Gene Ontology (GO)? Our study provides an investigation of the functional organization of five coexpression networks using data from three mammalian organisms. Our method calculates a probabilistic score between each gene and each Gene Ontology category that reflects coexpression enrichment of a GO module. For each GO category we use Receiver Operating Curves to assess whether these probabilistic scores reflect GBA. This methodology applied to five different coexpression networks demonstrates that the signature of guilt-by-association is ubiquitous and reproducible and that the GBA heuristic is broadly applicable across the population of nine hundred Gene Ontology categories. We also demonstrate the existence of highly reproducible patterns of coexpression between some pairs of GO categories. CONCLUSION: We conclude that GBA has universal value and that transcriptional control may be more modular than previously realized. Our analyses also suggest that methodologies combining coexpression measurements across multiple genes in a biologically-defined module can aid in characterizing gene function or in characterizing whether pairs of functions operate together.
Cecily J. Wolfe, Isaac S. Kohane, Atul J. Butte
BMC Bioinform.2
2005 Research Paper: Parents as Partners in Obtaining the Medication History
abstract
OBJECTIVE: Patient-centered information management may overcome barriers that impede high-quality, safe care in the emergency department (ED). The utility of parents' report of medication data via a multimedia, touch screen interface, the asthma kiosk, was investigated. Our specific aims were (1) to estimate the validity of parents' electronically entered medication history for asthma and (2) to compare the parents' kiosk entries regarding medications to the documentation of ED physicians and nurses. METHODS: We enrolled a cohort of parents to use the asthma kiosk and tested the validity of this communication channel for medication data specific to pediatric asthma. Parents' data provided via the kiosk during the ED encounter and the documentation of ED nurses and physicians were compared with a telephone-based interview with the parent after discharge that reviewed all asthma-specific medications physically present in the home. Treating clinicians in the ED were blinded to the parents' kiosk entries. RESULTS: Sixty-six parents were enrolled and 49 of 66 (74.2%) completed the gold standard interview. When analyzed at the level of individual medications, the validity of parental report was 81% for medication name, 79% for route of delivery, 66% for the form of the medication, and 60% for dose. Parents' report improved on the validity of documentation by physicians across all medication details save for medication name. Parents' report was more valid than nursing documentation at triage for all medication details. CONCLUSION: Parents can provide an independent source of medication data that improves on current documentation for key variables that impact quality and safety in emergency asthma care.
Stephen C. Porter, Isaac S. Kohane, Donald A. Goldmann
J. Am. Medical Informatics Assoc.2
2005 Position Paper: Wireless Technology Infrastructures for Authentication of Patients: PKI that Rings
abstract
As the public interest in consumer-driven electronic health care applications rises, so do concerns about the privacy and security of these applications. Achieving a balance between providing the necessary security while promoting user acceptance is a major obstacle in large-scale deployment of applications such as personal health records (PHRs). Robust and reliable forms of authentication are needed for PHRs, as the record will often contain sensitive and protected health information, including the patient's own annotations. Since the health care industry per se is unlikely to succeed at single-handedly developing and deploying a large scale, national authentication infrastructure, it makes sense to leverage existing hardware, software, and networks. This report proposes a new model for authentication of users to health care information applications, leveraging wireless mobile devices. Cell phones are widely distributed, have high user acceptance, and offer advanced security protocols. The authors propose harnessing this technology for the strong authentication of individuals by creating a registration authority and an authentication service, and examine the problems and promise of such a system.
Ulrich Sax, Isaac S. Kohane, Kenneth D. Mandl
J. Am. Medical Informatics Assoc.2
2005 Model Formulation: The PING Personally Controlled Electronic Medical Record System: Technical Architecture
abstract
Despite progress in creating standardized clinical data models and interapplication protocols, the goal of creating a lifelong health care record remains mired in the pragmatics of interinstitutional competition, concerns about privacy and unnecessary disclosure, and the lack of a nationwide system for authenticating and authorizing access to medical information. The authors describe the architecture of a personally controlled health care record system, PING, that is not institutionally bound, is a free and open source, and meets the policy requirements that the authors have previously identified for health care delivery and population-wide research.
William W. Simons, Kenneth D. Mandl, Isaac S. Kohane
J. Am. Medical Informatics Assoc.3
2004 Quantifying the relationship between co-expression, co-regulation and gene function
abstract
BACKGROUND: It is thought that genes with similar patterns of mRNA expression and genes with similar functions are likely to be regulated via the same mechanisms. It has been difficult to quantitatively test these hypotheses on a large scale because there has been no general way of determining whether genes share a common regulatory mechanism. Here we use data from a recent genome wide binding analysis in combination with mRNA expression data and existing functional annotations to quantify the likelihood that genes with varying degrees of similarity in mRNA expression profile or function will be bound by a common transcription factor. RESULTS: Genes with strongly correlated mRNA expression profiles are more likely to have their promoter regions bound by a common transcription factor. This effect is present only at relatively high levels of expression similarity. In order for two genes to have a greater than 50% chance of sharing a common transcription factor binder, the correlation between their expression profiles (across the 611 microarrays used in our study) must be greater than 0.84. Genes with similar functional annotations are also more likely to be bound by a common transcription factor. Combining mRNA expression data with functional annotation results in a better predictive model than using either data source alone. CONCLUSIONS: We demonstrate how mRNA expression data and functional annotations can be used together to estimate the probability that genes share a common regulatory mechanism. Existing microarray data and known functional annotations are sufficient to identify only a relatively small percentage of co-regulated genes.
Dominic J. Allocco, Isaac S. Kohane, Atul J. Butte
BMC Bioinform.2
2004 A SNP-centric database for the investigation of the human genome
abstract
BACKGROUND: Single Nucleotide Polymorphisms (SNPs) are an increasingly important tool for genetic and biomedical research. Although current genomic databases contain information on several million SNPs and are growing at a very fast rate, the true value of a SNP in this context is a function of the quality of the annotations that characterize it. Retrieving and analyzing such data for a large number of SNPs often represents a major bottleneck in the design of large-scale association studies. DESCRIPTION: SNPper is a web-based application designed to facilitate the retrieval and use of human SNPs for high-throughput research purposes. It provides a rich local database generated by combining SNP data with the Human Genome sequence and with several other data sources, and offers the user a variety of querying, visualization and data export tools. In this paper we describe the structure and organization of the SNPper database, we review the available data export and visualization options, and we describe how the architecture of SNPper and its specialized data structures support high-volume SNP analysis. CONCLUSIONS: The rich annotation database and the powerful data manipulation and presentation facilities it offers make SNPper a very useful online resource for SNP research. Its success proves the great need for integrated and interoperable resources in the field of computational biology, and shows how such systems may play a critical role in supporting the large-scale computational analysis of our genome.
Alberto Riva, Isaac S. Kohane
BMC Bioinform.2
2004 White Paper: Training the Next Generation of Informaticians: The Impact of "BISTI" and Bioinformatics - A Report from the American College of Medical Informatics
abstract
In 2002-2003, the American College of Medical Informatics (ACMI) undertook a study of the future of informatics training. This project capitalized on the rapidly expanding interest in the role of computation in basic biological research, well characterized in the National Institutes of Health (NIH) Biomedical Information Science and Technology Initiative (BISTI) report. The defining activity of the project was the three-day 2002 Annual Symposium of the College. A committee, comprised of the authors of this report, subsequently carried out activities, including interviews with a broader informatics and biological sciences constituency, collation and categorization of observations, and generation of recommendations. The committee viewed biomedical informatics as an interdisciplinary field, combining basic informational and computational sciences with application domains, including health care, biological research, and education. Consequently, effective training in informatics, viewed from a national perspective, should encompass four key elements: (1). curricula that integrate experiences in the computational sciences and application domains rather than just concatenating them; (2). diversity among trainees, with individualized, interdisciplinary cross-training allowing each trainee to develop key competencies that he or she does not initially possess; (3). direct immersion in research and development activities; and (4). exposure across the wide range of basic informational and computational sciences. Informatics training programs that implement these features, irrespective of their funding sources, will meet and exceed the challenges raised by the BISTI report, and optimally prepare their trainees for careers in a field that continues to evolve.
Charles P. Friedman, Russ B. Altman, Isaac S. Kohane, Kathleen A. McCormick, Perry L. Miller, Judy G. Ozbolt, Edward H. Shortliffe, Gary D. Stormo, M. Cleat Szczepaniak, David Tuck, Jeffrey J. Williamson
J. Am. Medical Informatics Assoc.3
2004 Application of Information Technology: The Asthma Kiosk: A Patient-centered Technology for Collaborative Decision Support in the Emergency Department
abstract
The authors report on the development and evaluation of a novel patient-centered technology that promotes capture of critical information necessary to drive guideline-based care for pediatric asthma. The design of this application, the asthma kiosk, addresses five critical issues for patient-centered technology that promotes guideline-based care: (1) a front-end mechanism for patient-driven data capture, (2) neutrality regarding patients' medical expertise and technical backgrounds, (3) granular capture of medication data directly from the patient, (4) formal algorithms linking patient-level semantics and asthma guidelines, and (5) output to both patients and clinical providers regarding best practice. The formative evaluation of the asthma kiosk demonstrates its ability to capture patient-specific data during real-time care in the emergency department (ED) with a mean completion time of 11 minutes. The asthma kiosk successfully links parents' data to guideline recommendations and identifies data critical to health improvements for asthmatic children that otherwise remains undocumented during ED-based care.
Stephen C. Porter, Zhaohui Cai, William Gribbons, Donald A. Goldmann, Isaac S. Kohane
J. Am. Medical Informatics Assoc.5
2003 Localization and Characterization of Mouse-Human Alignments Within the Human Genome. Does Evolutionary Conservation Suggest Functional Importance?
Sunil Saluja, Isaac S. Kohane
AMIA2
2003 PGAGENE: integrating quantitative gene-specific results from the NHLBI Programs for Genomic Applications
abstract
Abstract Summary: PGAGENE is a web-based gene-specific genomic data search engine, which allows users to search over 5.9 million pieces of collective genetic and genomic data from the NHLBI supported Programs for Genomic Applications. This data includes microarray measurements, SNPs, and mutations, and data may be found using symbols, parts of gene names or products, Affymetrix probe IDs, GenBank accession numbers, UniGene IDs, dbSNP IDs, and others. The PGAGENE indexing agent periodically maps all publicly available gene-specific PGA data onto LocusLink using dynamically generated cross-referencing tables. Availability: http://pgagene.chip.org Contact: [email protected] * To whom correspondence should be addressed.
Kyungjoon Lee, Isaac S. Kohane, Atul J. Butte
Bioinform.2
2003 Reproducibility of gene expression across generations of Affymetrix microarrays
abstract
BACKGROUND: The development of large-scale gene expression profiling technologies is rapidly changing the norms of biological investigation. But the rapid pace of change itself presents challenges. Commercial microarrays are regularly modified to incorporate new genes and improved target sequences. Although the ability to compare datasets across generations is crucial for any long-term research project, to date no means to allow such comparisons have been developed. In this study the reproducibility of gene expression levels across two generations of Affymetrix GeneChips (HuGeneFL and HG-U95A) was measured. RESULTS: Correlation coefficients were computed for gene expression values across chip generations based on different measures of similarity. Comparing the absolute calls assigned to the individual probe sets across the generations found them to be largely unchanged. CONCLUSION: We show that experimental replicates are highly reproducible, but that reproducibility across generations depends on the degree of similarity of the probe sets and the expression level of the corresponding transcript.
Ashish Nimgaonkar, Despina Sanoudou, Atul J. Butte, Judith N. Haslett, Louis M. Kunkel, Alan H. Beggs, Isaac S. Kohane
BMC Bioinform.7
2003 Editorial: Methods in Functional Genomics
Paola Sebastiani, Isaac S. Kohane, Marco Ramoni
Mach. Learn.2
2002 Computerized reminders to physicians in the emergency department: a web-based system to report late-arriving abnormal laboratory results
Zhaohui Cai, Isaac S. Kohane, Gary R. Fleisher, David S. Greenes
AMIA2
2002 Maximizing Data Portability in Patient-controlled Longitudinal Medical Records
Matvey Palchuk, Eric C. Pan, Isaac S. Kohane
AMIA3
2002 Patient-controlled Pediatric Immunization Records
Eric C. Pan, Isaac S. Kohane, Kenneth D. Mandl
AMIA2
2002 Accessing genomic data through XML-based remote procedure calls
Alberto Riva, Isaac S. Kohane
AMIA2
2002 An unsupervised self-optimizing gene clustering algorithm
Asher D. Schachter, Isaac S. Kohane
AMIA2
2002 CHIP TUNER: a web tool for evidence-based noise reduction in gene discovery
Christine L. Tsien, T. A. Libermann, A. Kho, Isaac S. Kohane
AMIA5
2002 Linking gene expression data with patient survival times using partial least squares
abstract
There is an increasing need to link the large amount of genotypic data, gathered using microarrays for example, with various phenotypic data from patients. The classification problem in which gene expression data serve as predictors and a class label phenotype as the binary outcome variable has been examined extensively, but there has been less emphasis in dealing with other types of phenotypic data. In particular, patient survival times with censoring are often not used directly as a response variable due to the complications that arise from censoring. We show that the issues involving censored data can be circumvented by reformulating the problem as a standard Poisson regression problem. The procedure for solving the transformed problem is a combination of two approaches: partial least squares, a regression technique that is especially effective when there is severe collinearity due to a large number of predictors, and generalized linear regression, which extends standard linear regression to deal with various types of response variables. The linear combinations of the original variables identified by the method are highly correlated with the patient survival times and at the same time account for the variability in the covariates. The algorithm is fast, as it does not involve any matrix decompositions in the iterations. We apply our method to data sets from lung carcinoma and diffuse large B-cell lymphoma studies to verify its effectiveness.
Peter J. Park, Isaac S. Kohane
ISMB3
2002 Analysis of matched mRNA measurements from two different microarray technologies
abstract
MOTIVATION: [corrected] The existence of several technologies for measuring gene expression makes the question of cross-technology agreement of measurements an important issue. Cross-platform utilization of data from different technologies has the potential to reduce the need to duplicate experiments but requires corresponding measurements to be comparable. METHODS: A comparison of mRNA measurements of 2895 sequence-matched genes in 56 cell lines from the standard panel of 60 cancer cell lines from the National Cancer Institute (NCI 60) was carried out by calculating correlation between matched measurements and calculating concordance between cluster from two high-throughput DNA microarray technologies, Stanford type cDNA microarrays and Affymetrix oligonucleotide microarrays. RESULTS: In general, corresponding measurements from the two platforms showed poor correlation. Clusters of genes and cell lines were discordant between the two technologies, suggesting that relative intra-technology relationships were not preserved. GC-content, sequence length, average signal intensity, and an estimator of cross-hybridization were found to be associated with the degree of correlation. This suggests gene-specific, or more correctly probe-specific, factors influencing measurements differently in the two platforms, implying a poor prognosis for a broad utilization of gene expression measurements across platforms.
Winston Patrick Kuo, Tor-Kristian Jenssen, Atul J. Butte, Lucila Ohno-Machado, Isaac S. Kohane
Bioinform.5
2002 Comparing expression profiles of genes with similar promoter regions
abstract
MOTIVATION: Gene regulatory elements are often predicted by seeking common sequences in the promoter regions of genes that are clustered together based on their expression profiles. We consider the problem in the opposite direction: we seek to find the genes that have similar promoter regions and determine the extent to which these genes have similar expression profiles. RESULTS: We use the data sets from experiments on Saccharomyces cerevisiae. Our similarity measure for the promoter regions is based on the set of common mapped or putative transcription factor binding sites and other regulatory elements in the upstream region of the genes, as contained in the Saccharomyces cerevisiae Promoter Database. We pair up the genes with high similarity scores and compare their expression levels in time-course experiment data. We find that genes with similar promoter regions on the average have significantly higher correlation, but it can vary widely depending on the genes. This confirms that the presence of similar regulatory elements often does not correspond to similarity in expression profiles and indicates that finding transcription factor binding sites or other regulatory elements starting with the expression patterns may be limited in many cases. Regardless of the correlation, the degree to which the profiles agree under different experimental conditions can be examined to derive hypotheses concerning the role of common regulatory elements. Overall, we find that considering the relationship between the promoter regions and the expression profiles starting with the regulatory elements is a difficult but useful process that can provide valuable insights.
Peter J. Park, Atul J. Butte, Isaac S. Kohane
Bioinform.3
2002 SNPper: retrieval and analysis of human SNPs
abstract
Abstract Motivation: Single Nucleotide Polymorphisms (SNPs) are an increasingly important tool for the study of the human genome. SNPs can be used as markers to create high-density genetic maps, as causal candidates for diseases, or to reconstruct the history of our genome. SNP-based studies rely on the availability of large numbers of validated, high-frequency SNPs whose position on the chromosomes is known with precision. Although large collections of SNPs exist in public databases, researchers need tools to effectively retrieve and manipulate them. Results: We describe the implementation and usage of SNPper, a web-based application to automate the tasks of extracting SNPs from public databases, analyzing them and exporting them in formats suitable for subsequent use. Our application is oriented toward the needs of candidate-gene, whole-genome and fine-mapping studies, and provides several flexible ways to present and export the data. The application has been publicly available for over a year, and has received positive user feedback and high usage levels. Availability: SNPper is freely available at http://snpper.chip.org/. Registration is optional and provides access to some advanced features. Contact: [email protected] * To whom correspondence should be addressed.
Alberto Riva, Isaac S. Kohane
Bioinform.2
2002 Brief Review: The Contributions of Biomedical Informatics to the Fight Against Bioterrorism
abstract
A comprehensive and timely response to current and future bioterrorist attacks requires a data acquisition, threat detection, and response infrastructure with unprecedented scope in time and space. Fortunately, biomedical informaticians have developed and implemented architectures, methodologies, and tools at the local and the regional levels that can be immediately pressed into service for the protection of our populations from these attacks. These unique contributions of the discipline of biomedical informatics are reviewed here.
Isaac S. Kohane
J. Am. Medical Informatics Assoc.1
2002 Visualization and evaluation of clusters for exploratory analysis of gene expression data
Ju-Han Kim, Isaac S. Kohane, Lucila Ohno-Machado
J. Biomed. Informatics2
2001 A web-based tool to retrieve human genome polymorphisms from public databases
Alberto Riva, Isaac S. Kohane
AMIA2
2001 Building ICU artifact detection models with more data in less time
Christine L. Tsien, Isaac S. Kohane, Neil McIntosh
AMIA2
2001 Evolution strategy applied to global optimization of clusters in gene expression data of DNA microarrays
abstract
Cluster analysis is the most important method for analyzing large-scale gene expression patterns. The matrix representation of microarray data and its successive 'optimal' incisional hyperplanes that create top-down hierarchical tree are a useful platform for developing optimization algorithms to determine the 'optimal' clusters from a pairwise proximity matrix which represents completely connected and weighted graph. Evolution strategy is applied to determine the 'globally optimal' incisional hyperplanes to construct hierarchical tree structure and tested with Fisher's iris and Golub's leukemia data sets. The results were compared with those of bottom-up hierarchical clustering, K-means and SOMs (Self-Organizing Maps) algorithms with promising results.
Kwonmoo Lee, Ju Han Kim, Tae Su Chung, Byoung-Sun Moon, Hoseung Lee, Isaac S. Kohane
CEC6
2001 Comparing the Similarity of Time-Series Gene Expression Using Signal Processing Metrics
Atul J. Butte, Ling Bao, Ben Y. Reis, Timothy W. Watkins, Isaac S. Kohane
J. Biomed. Informatics5
2001 Extracting Knowledge from Dynamics in Gene Expression
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane
J. Biomed. Informatics3
2001 Reply
Ben Y. Reis, Atul J. Butte, Isaac S. Kohane
J. Biomed. Informatics3
2000 Enrolling patients into clinical trials faster using RealTime Recuiting
Atul J. Butte, David A. Weinstein, Isaac S. Kohane
AMIA3
2000 The new peer review
Isaac S. Kohane, Russ B. Altman
AMIA1
2000 A Distributed, Secure File System For Personal Medical Records
Kenneth D. Mandl, Alberto Riva, Isaac S. Kohane
AMIA3
2000 Glucoweb: a case study of secure, remote biomonitoring and communication
Daniel J. Nigrin, Isaac S. Kohane
AMIA2
2000 Collating of a Distributed XML-based Medical Records into a Relational Database
Do Hoon Oh, Alberto Riva, Kenneth D. Mandl, Isaac S. Kohane
AMIA4
2000 First Steps Towards Implementing an International Training Program in Medical Informatics: The Brazil/USA Project
Lucila Ohno-Machado, Aziz A. Boxwala, Hamish S. F. Fraser, Robert A. Greenes, Isaac S. Kohane, Heimar F. Marin, Eduardo P. Marques, Eduardo Massad, Beatriz H. S. C. Rocha, Roberto A. Rocha, Laura M. Smeaton, Peter Szolovits
AMIA5
2000 Development of a Parent-completed Electronic Interview to Assess Dehydration
Stephen C. Porter, Gary R. Fleisher, Isaac S. Kohane, Kenneth D. Mandl
AMIA3
2000 Multiple signal integration by decision tree induction to detect artifacts in the neonatal intensive care unit
Christine L. Tsien, Isaac S. Kohane, Neil McIntosh
Artif. Intell. Medicine2
2000 Bioinformatics and Clinical Informatics: The Imperative to Collaborate
abstract
In this issue, Perry Miller and Russ Altman review the experiences at Yale and Stanford that have led to a convergence and cross-pollination between clinical informatics and bioinformatics at those institutions. A related convergence is revealed by a MEDLINE search for the string “informatics” in the last five months. Of 346 publications, 175 were in the area of genomics and not clinical applications. Concurrently, there has been much discussion within the informatics community about the dual nature of the research agenda (and, not coincidentally, the funding opportunities) as it pertains to clinical applications and fundamental biological research.1,2 Informal discussions with investigators in bioinformatics and clinical informatics, however, are tinged with concern that these two disciplines in biomedical informatics will diverge or at least that the two investigator communities are not collaborating sufficiently. It may, therefore, be timely to briefly review several major categories in which these two strains of biomedical informatics share common methodological and policy challenges. Moreover, as suggested by this overview, the success of bioinformatics and clinical informatics will depend on joint successes in resolving their mutual challenges. The categories outlined here are by no means intended to exhaustively cover the areas of commonality but are intended to provide a useful reference point for discussions on this topic of increasing relevance to the readership of JAMIA. In less than a decade, the Human Genome project (HGP)3 has generated a large amount of biological data that is likely eventually to lead to a qualitative change in the way in which clinical medicine (diagnostics, prognostics, and therapeutics) is practiced. A central intellectual and technologic asset to this effort has been GenBank4 and related genomic and protein databases (e.g., the SWISS-PROT,5 Exon-Intron,6 and IMGT databases7). Their standardized data models have allowed research laboratories throughout the world to rapidly populate them with the very latest information. In turn, these databases are freely available throughout the world via the Internet and have seeded, accelerated, and inspired thousands of research projects. In contrast, there are few, if any, consequential shared national clinical databases. Specifically, patient data in one information system can only rarely be transferred to another to expedite patient care. This, despite decades of research and development of clinical record systems. This marked contrast is deceptive. The HGP has benefited from the elegant simplicity of the genetic code. In essence, at the level of primary structure, the genetic information coded by any organism is simply a sequence of characters drawn from a very limited alphabet. Consequently, there are only a very few items that GenBank requires be submitted for an entry to be a valid (and useful) component of its database. The clinical care of human beings is far more complex, requiring at the minimum a detailed record of the history of multiple clinical interventions and outcomes, relevant life history, and clinical measurements that span several modalities, from serum chemistry to brain imaging. It is not surprising that the data model required to capture all this information is extremely complex, as is evidenced by the Health Level 7 Reference Information Model.8,9 It is a remarkable tribute to the persistence of the individuals involved in these standardization efforts, that they have been able to arrive at a reasonably adequate standardized representation of not only the many descriptors but much of the process and business of clinical care. As the HGP moves from the acquisition of raw genomics data to the biological function of the discovered genes and their clinical importance, the bioinformatics community will have to address very similar complexities. That is, the clinical annotation of genomic data sets, particularly for human beings, will essentially provide the equivalent, if not identical, challenge of the creation of a comprehensive medical record. Even prior to encompassing the entirety of clinical annotation, the genomics community has faltered in developing shared and standardized data models where the simplicity of the genome no longer dominates. For example, there are several competing technologies for the massively parallel measurement of gene expression using microarrays. Some of these arrays use two probes per gene and are constructed using robotic spotting techniques. Others are constructed with oligonucleotides using photolithographic techniques.10 Although all these techniques measure gene expression, a widely adopted standard to represent the results across all microarray technologies has yet to emerge. The GATC11 proposal, for example, is a possible candidate for such a data model, but its usage is currently spotty and controversial. To clinical informaticians, this will be all too reminiscent of the challenge of creating a shared data model for laboratory results. The lack of widely accepted standardized vocabularies for clinical care has greatly hampered the development of automated decision support tools and clinical research databases. The impossibility of guaranteeing that a serum sodium or systolic blood pressure has the same code or term throughout our hospital system is troublesome. Fortunately, several efforts in the private and public realm (e.g., LOINC12) are addressing this issue. The National Library of Medicine has invested large resources to enable these different vocabularies to be interoperable, at least at a basic level.13 The same problems are not unknown in bioinformatics. Even at this early stage of the HGP, DNA sequences that were previously not known to be part of the same gene have different names and are joined in only some databases (with varying levels of confidence). As the HGP ventures into diverse areas of bioscience (as well as into the clinical area), vocabulary issues are also important. Indeed, the lack of a standardized vocabulary already arises in genomics as well, in annotations. For example, despite the fact that the basic data element of GenBank is the sequence (which has an easily standardized representation), there are diverse annotations that are very nonstandardized right now. A recent report by the Institute of Medicine14 highlights the immense mortality and morbidity due to medical errors. Clinical informaticians (e.g., Bates et al.15 and Kuperman et al.16,17) have been instrumental in demonstrating how automated systems can be used to reduce this error rate. These industrial processing and quality improvement techniques are not without relevance to the HGP. It is well known that the mouse genome database has been contaminated with entries of the rat genome and that the specification of 5′ to 3′ polarity of a gene sequence has been found to be inverted.18–20 And these are only some of the known errors in a very large effort. These sources of error can be reduced, as they have in many industries, by the application of increased process automation and automated interception of human error before it becomes consequential. The architectures of clinical order entry systems, designed for complex clinical enterprises to prevent erroneous and dangerous clinical behavior, can inform the design of genome sequencing and expression profiling processes to prevent the kind of errors we are already finding in genomic databases. It is well known that the noise in clinical measurements leads to erroneous decision making. The archetypal example is in the intensive care unit, where multiple physiologic monitors each has its own alarm module. Because of the noisy nature of the biological signals that are monitored, the alarms are ignored or switched off because of their high false-positive rate.21 The noisy nature of the monitored signals thus has a significant impact on the provision of care and the decision-making ability of care providers who are working under conditions of uncertainty and data overload. Similar noise considerations arise in genomics. For example, with gene microarrays, we can measure the expression of tens of thousands of genes at a time. There are several sources of noise in these measurements: within a microarray, across microarrays, and from the intrinsic variability of the biological systems being measured. Yet in 1999, several reports, which appeared in scientific journals of the first rank, included changes in expression so small as to be indistinguishable from noise. Such changes are, in essence, a false-positive result. These false positives are potentially extremely costly. A biological researcher might decide to invest several months investigating a gene's regulation because a microarray experiment showed it to be increased or decreased under a particular set of conditions. In clinical informatics there is a rich literature of the techniques that can be used to identify false positives and reduce noise (e.g., filtering, signal fusion22–26). Many of these techniques are transferable to the genomic domain. In 1997, the Institute of Medicine27 reported significant lacunae in both technology and policy in protecting confidential patient data. Among the problems of greatest concern that were emphasized by this report were the relatively unrestricted access by third parties to these data for secondary uses and the inadequacy of the anonymization process (in both practice and theory28,29). Subsequently, the clinical informatics community has developed several model confidentiality policies30 and cryptographic identification systems.31 As the fruits of the HGP are translated first into clinical research protocols and then into clinical practice, personally identifiable genomic data will find their way into some form of information system. The challenges posed to the security and privacy of such data will dwarf any encountered to date with conventional clinical data. The reasons are twofold: First, genomic information is likely to be much more predictive of current and future health status than most clinical measurements. Second, with very few exceptions, an individual's genome is uniquely identifying. This identifiability is much more reliable, persistent, and specific than typically cited identifiers, including a person's name, social security number, date of birth, and address. At the very least, the architects of information systems storing genetic data should learn from all the mistakes of and designs developed for the security architectures and privacy policies of conventional clinical information systems. Conversely, the extreme concerns posed by the storage of personal genetic data is likely to generate new policies and security architectures that will enhance the confidentiality of clinical information systems. Moreover, when personal genetic data becomes incorporated into routine medical practice, the safeguards for the confidentiality of the medical record will be crucial to the confidentiality of the genetic data referenced there. “Getting the data in” has often been cited32 by authorities in clinical informatics as being among the most difficult challenges in successfully deploying clinical information systems. In particular, the costs of acquiring detailed and structured data from the clinical care process have been daunting. Voice and handwriting recognition information systems have not been broadly adopted, because of a variety of performance and usability issues. The cost and practicability issues will continue to present obstacles to clinical information system utility and deployment until better solutions are arrived at. In contrast, the HGP has managed to achieve significant economies of scale in sequencing technology. Gene microarrays alone have dropped in cost by a factor of two in just the last year. Here again, once genomic investigators attempt to bridge the gulf from purely genomic data sets to phenotypically (i.e., clinically) annotated data sets, they will be confronted with the same challenges of clinically oriented, codified data acquisition. The questions of which user interfaces are the most cost efficient, reliable, and generalizable to multiple clinical domains are among the implementation and design challenges that they will face. Although they have yet to arrive at definitively successful answers, clinical informaticians have already completed several decades worth of engineering and ethnographic studies33 addressing the very same questions. The first rough draft of the human genome was reported to have been completed in May of this year.34 It is likely that a complete, high-quality human DNA reference sequence will be available by 2003. Yet the function of the vast majority of the genes in the human genome will be unknown. The minority of genes with documented function are likely to have many more functions and interactions that are unknown. Consequently, one of the primary applications of information technologies in genomics is the application of machine learning techniques to determine how genes are functionally interdependent and how these interdependencies are reflected in the biological and clinical behavior of the system in which they operate.3,35 Many of these machine learning techniques were previously applied to the task of extracting knowledge from clinical databases, and some were even developed first in the clinical domain.36–44 To be sure, the genomic era has challenged these machine learning techniques to the extreme, because of the high dimensionality of data sets (e.g., tens of thousands of measurements per experiment) and the relatively few cases and experiments from which investigators are attempting to glean knowledge. Without being exhaustive, this brief review suggests the multiple points of commonality between the genomic and clinical strands of the biomedical informatics research agenda. It also suggests that the training of investigators in informatics should include a set of core competencies that at least cover these common points. In this fashion, the joint research agenda might be well served to the mutual benefit of biomedical science and clinical care. The Stanford Medical Informatics educational program, described in this issue, illustrates this benefit.
Isaac S. Kohane
J. Am. Medical Informatics Assoc.1
2000 Application of Information Technology: Temporal Expressiveness in Querying a Time-stamp - based Clinical Database
abstract
Most health care databases include time-stamped instant data as the only temporal representation of patient information. Many previous efforts have attempted to provide frameworks in which medical databases could be queried in relation to time. These, however, have required either a sophisticated database representation of time, including time intervals, or a time-stamp-based database coupled with a nonstandard temporal query language. In this work, the authors demonstrate how their previously described data retrieval application, DXtractor, can be used as a database querying application with expressive power close to that of temporal databases and temporal query languages, using only standard SQL and existing time-stamp-based repositories. DXtractor provides the ability to compose temporal queries through an interface that is understood by nonprogramming medical personnel. Not all temporal constructs are easily implemented using this framework; nonetheless, DXtractor's temporal capabilities provide a significant improvement in the temporal expressivity accessible to clinicians using standard time-stamped clinical databases.
Daniel J. Nigrin, Isaac S. Kohane
J. Am. Medical Informatics Assoc.2
1999 Unsupervised knowledge discovery in medical databases using relevance networks
Atul J. Butte, Isaac S. Kohane
AMIA2
1999 Artifact detection in cardiovascular time series monitoring data from preterm infants
Isaac S. Kohane, Neil McIntosh
AMIA2
1999 Healthconnect: clinical grade patient-physician communication
Kenneth D. Mandl, Isaac S. Kohane
AMIA2
1999 Scaling a data retrieval and mining application to the enterprise-wide level
Daniel J. Nigrin, Isaac S. Kohane
AMIA2
1998 ParentLink: A Web-Based Communications Tool for Parents and Pediatricians
Dilek A. Bishku, Charles J. Homer, Kenneth D. Mandl, Isaac S. Kohane
AMIA4
1998 HealthConnect: A Structured Communication System for Health Management
Karen L. Bradshaw, Kenneth D. Mandl, Isaac S. Kohane
AMIA3
1998 An Alert System for ED Laboratory Test Results
Hongmei Dong, David S. Greenes, Gary R. Fleisher, Isaac S. Kohane
AMIA4
1998 Health information identification and de-identification toolkit
Isaac S. Kohane, Hongmei Dong, Peter Szolovits
AMIA1
1998 Social equity and access to the World Wide Web and E-mail: implications for design and implementation of medical applications
Kenneth D. Mandl, S. B. Katz, Isaac S. Kohane
AMIA3
1998 Data mining by clinicians
Daniel J. Nigrin, Isaac S. Kohane
AMIA2
1998 Parental Input into the Emergency Department Medical Record
Stephen C. Porter, Mary Silvia, Gary R. Fleisher, Isaac S. Kohane, Kenneth D. Mandl
AMIA4
1998 Methodology for Computer-Assisted Prediction of Growth Hormone Deficiency in Children
Christine L. Tsien, Daniel J. Nigrin, Isaac S. Kohane
AMIA3
1998 Linking multiple heterogeneous data sources to practice guidelines
F. J. van Wingerde, Oren Harary, Kenneth D. Mandl, Susanne Salem-Schatz, Charles J. Homer, Isaac S. Kohane
AMIA7
1998 LRTree: A Hybrid technique for Classifying Myocardial Infarction Data Containing Unknown Attribute Values
Christine L. Tsien, Hamish S. F. Fraser, Isaac S. Kohane
PAKDD3
1997 Confidentiality of Medical Records in the W3-EMRS Project
David M. Rind, Isaac S. Kohane, Peter Szolovits, Charles Safran, Henry C. Chueh, G. Octo Barnett
AMIA2
1997 A Java-based multi-institutional medical information retrieval system
F. J. van Wingerde, Karen L. Bradshaw, Peter Szolovits, Isaac S. Kohane
AMIA5
1996 Managing temporal worlds for medical trend diagnosis
Ira J. Haimowitz, Isaac S. Kohane
Artif. Intell. Medicine2
1996 Application of Technology: Building National Electronic Medical Record Systems via the World Wide Web
abstract
Electronic medical record systems (EMRSs) currently do not lend themselves easily to cross-institutional clinical care and research. Unique system designs coupled with a lack of standards have led to this difficulty. The authors have designed a preliminary EMRS architecture (W3-EMRS) that exploits the multiplatform, multiprotocol, client-server technology of the World Wide Web. The architecture abstracts the clinical information model and the visual presentation away from the underlying EMRS. As a result, computation upon data elements of the EMRS and their presentation are no longer tied to the underlying EMRS structures. The architecture is intended to enable implementation of programs that provide uniform access to multiple, heterogeneous legacy EMRSs. The authors have implemented an initial prototype of W3-EMRS that accesses the database of the Boston Children's Hospital Clinician's Workstation.
Isaac S. Kohane, Philip Greenspun, James C. Fackler, Christopher Cimino, Peter Szolovits
J. Am. Medical Informatics Assoc.1
1995 Clinical monitoring using regression-based trend templates
Ira J. Haimowitz, P. P. Le, Isaac S. Kohane
Artif. Intell. Medicine3
1995 Editorial
Isaac S. Kohane
Artif. Intell. Medicine1
1994 Policy Forum: Against Simple Universal Health-Care Identifiers
abstract
Peter Szolovits, PhD, Isaac Kohane, MD, PhD; Against Simple Universal Health-care Identifiers, Journal of the American Medical Informatics Association, Volume 1
Peter Szolovits, Isaac S. Kohane
J. Am. Medical Informatics Assoc.2
1993 An Epistemology for Clinically Significant Trends
Ira J. Haimowitz, Isaac S. Kohane
AAAI2
1993 Integration of intermittent clinical data with continuous data from bedside monitors
abstract
Intermittent data from hospital information systems and continuous data from beside monitors are rarely integrated. This integration is crucial to providing physicians with comprehensive views of their patients' clinical state. It also can provide a broader clinical context for intelligent monitoring programs and thereby increase the sensitivity and specificity of their alarms. We describe an implemented application programming interface (API) that provides an abstracted query interface to the hospital's central, relational database and to data extracted from the monitors. The API also provides several data-reduction operators to help manage the large volume of data generated. We describe an electronic flowsheet that was built using the API.>
James C. Fackler, Isaac S. Kohane
CBMS2
1993 Automated Trend Detection with Alternate Temporal Hypotheses
Ira J. Haimowitz, Isaac S. Kohane
IJCAI2