VLDB 2026 Research / reviewers in the wild / expert
Shawn N. Murphy
dblp:68/7183
· DBLP profile ↗
129ranked-venue papers
20as first author
30since 2021 · last 2025
0000-0002-1905-8806ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 128 · 20 first-author · 30 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comparing patient-reported symptoms and structured clinician documentation in electronic health recordsabstractOBJECTIVES: Real-world data (RWD) analyses primarily rely on structured clinical documentation collected through routine clinical care or driven by medical billing requirements. Patient-reported outcome measures (PROMs), integrated into electronic health records (EHRs), are an additional data source that could offer valuable insights into a patient's perspective and contribute to a more comprehensive understanding of health outcomes in RWD studies. This study aims to characterize agreement between PROMs symptoms and structured clinical documentation of these symptoms by clinicians in EHRs. MATERIALS AND METHODS: A cross-sectional study of 913 244 adult primary care annual physical visits between January 1, 2019 and December 31, 2023. We compared differences in prevalence and agreement of patient-reported symptoms (PRS) and structured clinician documentation (CD) across 15 respiratory, gastrointestinal, cardiometabolic, and neuropsychiatric symptoms. RESULTS: Patient-reported symptom prevalence were significantly higher compared to CD across most symptoms including joint pain (33% PRS vs 12%), headaches (17% PRS vs 8.8% CD), and sleep disturbance (24% PRS vs 10% CD). Clinicians documented anxiety (11% PRS vs 23% CD) and depression (6.6% PRS vs 15.4% CD) symptoms using structured code at higher rates than patients reported them. Agreement between symptom self-report and clinician-documented structured codes was low to moderate (κ: 0.06-0.39). DISCUSSION: Primary care patients self-report symptoms up to ten times more frequently than clinicians document them with structured codes in the EHR. CONCLUSION: This work demonstrates the value and feasibility of incorporating PRSs in RWD studies to reduce misclassification and more holistically capture a patient's health. Victor M. Castro, Vivian S. Gainer, Danielle M. Crookes, Shawn N. Murphy, Justin Manjourides |
J. Am. Medical Informatics Assoc. | 4 |
| 2024 | Towards cross-application model-agnostic federated cohort discoveryabstractOBJECTIVES: To demonstrate that 2 popular cohort discovery tools, Leaf and the Shared Health Research Information Network (SHRINE), are readily interoperable. Specifically, we adapted Leaf to interoperate and function as a node in a federated data network that uses SHRINE and dynamically generate queries for heterogeneous data models. MATERIALS AND METHODS: SHRINE queries are designed to run on the Informatics for Integrating Biology & the Bedside (i2b2) data model. We created functionality in Leaf to interoperate with a SHRINE data network and dynamically translate SHRINE queries to other data models. We randomly selected 500 past queries from the SHRINE-based national Evolve to Next-Gen Accrual to Clinical Trials (ENACT) network for evaluation, and an additional 100 queries to refine and debug Leaf's translation functionality. We created a script for Leaf to convert the terms in the SHRINE queries into equivalent structured query language (SQL) concepts, which were then executed on 2 other data models. RESULTS AND DISCUSSION: 91.1% of the generated queries for non-i2b2 models returned counts within 5% (or ±5 patients for counts under 100) of i2b2, with 91.3% recall. Of the 8.9% of queries that exceeded the 5% margin, 77 of 89 (86.5%) were due to errors introduced by the Python script or the extract-transform-load process, which are easily fixed in a production deployment. The remaining errors were due to Leaf's translation function, which was later fixed. CONCLUSION: Our results support that cohort discovery applications such as Leaf and SHRINE can interoperate in federated data networks with heterogeneous data models. Nicholas J. Dobbins, Michele Morris, Eugene Sadhu, Douglas MacFadden, Marc-Danie Nazaire, William Simons, Griffin M. Weber, Shawn N. Murphy, Shyam Visweswaran |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | A broadly applicable approach to enrich electronic-health-record cohorts by identifying patients with complete data: a multisite evaluationabstractOBJECTIVE: Patients who receive most care within a single healthcare system (colloquially called a "loyalty cohort" since they typically return to the same providers) have mostly complete data within that organization's electronic health record (EHR). Loyalty cohorts have low data missingness, which can unintentionally bias research results. Using proxies of routine care and healthcare utilization metrics, we compute a per-patient score that identifies a loyalty cohort. MATERIALS AND METHODS: We implemented a computable program for the widely adopted i2b2 platform that identifies loyalty cohorts in EHRs based on a machine-learning model, which was previously validated using linked claims data. We developed a novel validation approach, which tests, using only EHR data, whether patients returned to the same healthcare system after the training period. We evaluated these tools at 3 institutions using data from 2017 to 2019. RESULTS: Loyalty cohort calculations to identify patients who returned during a 1-year follow-up yielded a mean area under the receiver operating characteristic curve of 0.77 using the original model and 0.80 after calibrating the model at individual sites. Factors such as multiple medications or visits contributed significantly at all sites. Screening tests' contributions (eg, colonoscopy) varied across sites, likely due to coding and population differences. DISCUSSION: This open-source implementation of a "loyalty score" algorithm had good predictive power. Enriching research cohorts by utilizing these low-missingness patients is a way to obtain the data completeness necessary for accurate causal analysis. CONCLUSION: i2b2 sites can use this approach to select cohorts with mostly complete EHR data. Jeffrey G. Klann, Darren W. Henderson, Michele Morris, Hossein Estiri, Griffin M. Weber, Shyam Visweswaran, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 7 |
| 2023 | Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes |
J. Biomed. Informatics | 24 |
| 2022 | A Deductive Data-Driven Pipeline Powered by MLHO for Post-Acute Sequelae of COVID-19 (PASC) Phenotyping
Arianna Dagliati, Zachary H. Strasser, Rebecca Mesa, Zahra Shakeri, Alaleh Azhir, Riccardo Bellazzi, Shawn N. Murphy, Hossein Estiri |
AMIA | 7 |
| 2022 | Informatics for Integrating Biology and the Bedside (i2b2) in 2022: Single Sign On and Synthetic Data
Jeffrey G. Klann, Michael Mendis, Kevin Bui, Griffin M. Weber, Diane Keogh, Shawn N. Murphy |
AMIA | 6 |
| 2022 | Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions
Jeffrey G. Klann, Zachary H. Strasser, Chris J. Kennedy, Meghan Hutch, John H. Holmes, Gabriel A. Brat, Shawn N. Murphy |
AMIA | 7 |
| 2022 | Research Patient Data Repositories: Perspectives from JAMIA Special Issue Editors on the Next Generation of Multi-Institutional Data Sharing
Genevieve B. Melton, Leslie Lenert, Michael J. Becich, Shawn N. Murphy, Thomas R. Campion Jr. |
AMIA | 4 |
| 2022 | Dynamic Reaction Picklist for Improving Allergy Reaction Documentation: A Usability Study
Heekyong Park, Sachin Vallamkonda, Diane L. Seger, Suzanne V. Blackley, Pam Garabedian, Foster R. Goss, Kimberly G. Blumenthal, David W. Bates, Shawn N. Murphy, Li Zhou 0007 |
AMIA | 10 |
| 2022 | The Informatics of RECOVER: Understanding the Post Acute Sequelae of SARS-CoV-2 Infection
Mark G. Weiner, L. Charles Bailey, Richard R. Moffitt, Shawn N. Murphy |
AMIA | 4 |
| 2022 | I2b2-etl: Python application for importing electronic health data into the informatics for integrating biology and the bedside platformabstractMOTIVATION: The i2b2 platform is used at major academic health institutions and research consortia for querying for electronic health data. However, a major obstacle for wider utilization of the platform is the complexity of data loading that entails a steep curve of learning the platform's complex data schemas. To address this problem, we have developed the i2b2-etl package that simplifies the data loading process, which will facilitate wider deployment and utilization of the platform. RESULTS: We have implemented i2b2-etl as a Python application that imports ontology and patient data using simplified input file schemas and provides inbuilt record number de-identification and data validation. We describe a real-world deployment of i2b2-etl for a population-management initiative at MassGeneral Brigham. AVAILABILITY AND IMPLEMENTATION: i2b2-etl is a free, open-source application implemented in Python available under the Mozilla 2 license. The application can be downloaded as compiled docker images. A live demo is available at https://i2b2clinical.org/demo-i2b2etl/ (username: demo, password: Etl@2021). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kavishwar B. Wagholikar, Layne Ainsworth, David Zelle, Kira Chaney, Michael Mendis, Jeffrey G. Klann, Alexander J. Blood, Angela Miller, Rupendra Chulyadyo, Michael Oates, William J. Gordon, Samuel J. Aronson, Benjamin M. Scirica, Shawn N. Murphy |
Bioinform. | 14 |
| 2022 | The Mass General Brigham Biobank Portal: an i2b2-based data repository linking disparate and high-dimensional patient data to support multimodal analyticsabstractOBJECTIVE: Integrating and harmonizing disparate patient data sources into one consolidated data portal enables researchers to conduct analysis efficiently and effectively. MATERIALS AND METHODS: We describe an implementation of Informatics for Integrating Biology and the Bedside (i2b2) to create the Mass General Brigham (MGB) Biobank Portal data repository. The repository integrates data from primary and curated data sources and is updated weekly. The data are made readily available to investigators in a data portal where they can easily construct and export customized datasets for analysis. RESULTS: As of July 2021, there are 125 645 consented patients enrolled in the MGB Biobank. 88 527 (70.5%) have a biospecimen, 55 121 (43.9%) have completed the health information survey, 43 552 (34.7%) have genomic data and 124 760 (99.3%) have EHR data. Twenty machine learning computed phenotypes are calculated on a weekly basis. There are currently 1220 active investigators who have run 58 793 patient queries and exported 10 257 analysis files. DISCUSSION: The Biobank Portal allows noninformatics researchers to conduct study feasibility by querying across many data sources and then extract data that are most useful to them for clinical studies. While institutions require substantial informatics resources to establish and maintain integrated data repositories, they yield significant research value to a wide range of investigators. CONCLUSION: The Biobank Portal and other patient data portals that integrate complex and simple datasets enable diverse research use cases. i2b2 tools to implement these registries and make the data interoperable are open source and freely available. Victor M. Castro, Vivian S. Gainer, Nich Wattanasin, Barbara Benoit, Andrew Cagan, Bhaswati Ghosh, Sergey Goryachev, Reeta Metta, Heekyong Park, Taowei David Wang, Michael Mendis, Martin Rees, Christopher Herrick, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 14 |
| 2022 | An objective framework for evaluating unrecognized bias in medical AI models predicting COVID-19 outcomesabstractOBJECTIVE: The increasing translation of artificial intelligence (AI)/machine learning (ML) models into clinical practice brings an increased risk of direct harm from modeling bias; however, bias remains incompletely measured in many medical AI applications. This article aims to provide a framework for objective evaluation of medical AI from multiple aspects, focusing on binary classification models. MATERIALS AND METHODS: Using data from over 56 000 Mass General Brigham (MGB) patients with confirmed severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), we evaluate unrecognized bias in 4 AI models developed during the early months of the pandemic in Boston, Massachusetts that predict risks of hospital admission, ICU admission, mechanical ventilation, and death after a SARS-CoV-2 infection purely based on their pre-infection longitudinal medical records. Models were evaluated both retrospectively and prospectively using model-level metrics of discrimination, accuracy, and reliability, and a novel individual-level metric for error. RESULTS: We found inconsistent instances of model-level bias in the prediction models. From an individual-level aspect, however, we found most all models performing with slightly higher error rates for older patients. DISCUSSION: While a model can be biased against certain protected groups (ie, perform worse) in certain tasks, it can be at the same time biased towards another protected group (ie, perform better). As such, current bias evaluation studies may lack a full depiction of the variable effects of a model on its subpopulations. CONCLUSION: Only a holistic evaluation, a diligent search for unrecognized bias, can provide enough information for an unbiased judgment of AI bias that can invigorate follow-up investigations on identifying the underlying roots of bias and ultimately make a change. Hossein Estiri, Zachary H. Strasser, Sina Rashidian, Jeffrey G. Klann, Kavishwar B. Wagholikar, Thomas H. McCoy Jr., Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | A framework for employing longitudinally collected multicenter electronic health records to stratify heterogeneous patient populations on disease historyabstractOBJECTIVE: To facilitate patient disease subset and risk factor identification by constructing a pipeline which is generalizable, provides easily interpretable results, and allows replication by overcoming electronic health records (EHRs) batch effects. MATERIAL AND METHODS: We used 1872 billing codes in EHRs of 102 880 patients from 12 healthcare systems. Using tools borrowed from single-cell omics, we mitigated center-specific batch effects and performed clustering to identify patients with highly similar medical history patterns across the various centers. Our visualization method (PheSpec) depicts the phenotypic profile of clusters, applies a novel filtering of noninformative codes (Ranked Scope Pervasion), and indicates the most distinguishing features. RESULTS: We observed 114 clinically meaningful profiles, for example, linking prostate hyperplasia with cancer and diabetes with cardiovascular problems and grouping pediatric developmental disorders. Our framework identified disease subsets, exemplified by 6 "other headache" clusters, where phenotypic profiles suggested different underlying mechanisms: migraine, convulsion, injury, eye problems, joint pain, and pituitary gland disorders. Phenotypic patterns replicated well, with high correlations of ≥0.75 to an average of 6 (2-8) of the 12 different cohorts, demonstrating the consistency with which our method discovers disease history profiles. DISCUSSION: Costly clinical research ventures should be based on solid hypotheses. We repurpose methods from single-cell omics to build these hypotheses from observational EHR data, distilling useful information from complex data. CONCLUSION: We establish a generalizable pipeline for the identification and replication of clinically meaningful (sub)phenotypes from widely available high-dimensional billing codes. This approach overcomes datatype problems and produces comprehensive visualizations of validation-ready phenotypes. Marc P. Maurits, Ilya Korsunsky, Soumya Raychaudhuri, Shawn N. Murphy, Jordan W. Smoller, Scott T. Weiss, Thomas W. J. Huizinga, Marcel J. T. Reinders, Elizabeth W. Karlson, Erik van den Akker 0001, Rachel Knevel |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Research data warehouse best practices: catalyzing national data sharing through informatics innovationabstractResearch Patient Data Repositories (RPDRs) have become essential infrastructure for traditional Clinical and Translational Science Award (CTSA) programs and increasingly for a wide range of research consortia and learning health system networks.1–5 Almost every institution with a CTSA or Clinical Translational Research (CTR) program (found in states with lower amounts of National Institutes of Health funding) hosts an RPDR for the benefit of affiliated researchers. These repositories aim to enable healthcare research based upon the patient populations they serve. Within the institution, RPDRs are valuable for a range of research activities. They are used to identify patients for clinical trial recruitment using privacy-preserving methods to search and extract specific cohorts of trial-eligible patients.6 They aid in developing and validating computable phenotypes that are increasingly important for accurately identifying patient cohorts in a reproducible fashion.7 RPDRs provide de-identified patient data for population health research and support a growing body of artificial intelligence to predict patient outcomes.8 Further, clinical studies can often be simulated using data from an RPDR.9 Beyond the institution, aggregates of de-identified datasets from multiple institutions linked with privacy-preserving hash codes provide an unprecedented opportunity to conduct population health research, perform comparative effectiveness analyses and apply artificial intelligence methods over large and diverse populations.10 The data contained within the RPDR vary across institutions, based on institutional strengths and weaknesses; the papers published in this issue reflect that variability (see Table 1). Data are commonly acquired from local electronic health records (EHRs) and other clinical information systems that capture information during clinical care. Data consist of diagnoses, problem lists, procedures, prescribed medications, laboratory exams, and many types of free-text reports. Overall, the benefits of the RPDR for accelerating translational research can be significant. For example, at Harvard, in 2006, between $94 and $136 million in annual research funding was linked to the use of data from the RPDR.11 Shawn N. Murphy, Shyam Visweswaran, Michael J. Becich, Thomas R. Campion Jr., Boyd M. Knosp, Genevieve B. Melton, Leslie Lenert |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Analytics to monitor local impact of the Protecting Access to Medicare Act's imaging clinical decision support requirementsabstractOBJECTIVE: This study aimed is to: (1) extend the Integrating the Biology and the Bedside (i2b2) data and application models to include medical imaging appropriate use criteria, enabling it to serve as a platform to monitor local impact of the Protecting Access to Medicare Act's (PAMA) imaging clinical decision support (CDS) requirements, and (2) validate the i2b2 extension using data from the Medicare Imaging Demonstration (MID) CDS implementation. MATERIALS AND METHODS: This study provided a reference implementation and assessed its validity and reliability using data from the MID, the federal government's predecessor to PAMA's imaging CDS program. The Star Schema was extended to describe the interactions of imaging ordering providers with the CDS. New ontologies were added to enable mapping medical imaging appropriateness data to i2b2 schema. z-Ratio for testing the significance of the difference between 2 independent proportions was utilized. RESULTS: The reference implementation used 26 327 orders for imaging examinations which were persisted to the modified i2b2 schema. As an illustration of the analytical capabilities of the Web Client, we report that 331/1192 or 28.1% of imaging orders were deemed appropriate by the CDS system at the end of the intervention period (September 2013), an increase from 162/1223 or 13.2% for the first month of the baseline period, December 2011 (P = .0212), consistent with previous studies. CONCLUSIONS: The i2b2 platform can be extended to monitor local impact of PAMA's appropriateness of imaging ordering CDS requirements. Vladimir I. Valtchinov, Shawn N. Murphy, Ronilda C. Lacson, Nikolay Ikonomov, Bingxue K. Zhai, Katherine P. Andriole, Justin F. Rousseau, Dick Hanson, Isaac S. Kohane, Ramin Khorasani |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai |
J. Biomed. Informatics | 34 |
| 2022 | Multiview Incomplete Knowledge Graph Integration with application to cross-institutional EHR data harmonizationabstractOBJECTIVE: The growing availability of electronic health records (EHR) data opens opportunities for integrative analysis of multi-institutional EHR to produce generalizable knowledge. A key barrier to such integrative analyses is the lack of semantic interoperability across different institutions due to coding differences. We propose a Multiview Incomplete Knowledge Graph Integration (MIKGI) algorithm to integrate information from multiple sources with partially overlapping EHR concept codes to enable translations between healthcare systems. METHODS: The MIKGI algorithm combines knowledge graph information from (i) embeddings trained from the co-occurrence patterns of medical codes within each EHR system and (ii) semantic embeddings of the textual strings of all medical codes obtained from the Self-Aligning Pretrained BERT (SAPBERT) algorithm. Due to the heterogeneity in the coding across healthcare systems, each EHR source provides partial coverage of the available codes. MIKGI synthesizes the incomplete knowledge graphs derived from these multi-source embeddings by minimizing a spherical loss function that combines the pairwise directional similarities of embeddings computed from all available sources. MIKGI outputs harmonized semantic embedding vectors for all EHR codes, which improves the quality of the embeddings and enables direct assessment of both similarity and relatedness between any pair of codes from multiple healthcare systems. RESULTS: With EHR co-occurrence data from Veteran Affairs (VA) healthcare and Mass General Brigham (MGB), MIKGI algorithm produces high quality embeddings for a variety of downstream tasks including detecting known similar or related entity pairs and mapping VA local codes to the relevant EHR codes used at MGB. Based on the cosine similarity of the MIKGI trained embeddings, the AUC was 0.918 for detecting similar entity pairs and 0.809 for detecting related pairs. For cross-institutional medical code mapping, the top 1 and top 5 accuracy were 91.0% and 97.5% when mapping medication codes at VA to RxNorm medication codes at MGB; 59.1% and 75.8% when mapping VA local laboratory codes to LOINC hierarchy. When trained with 500 labels, the lab code mapping attained top 1 and 5 accuracy at 77.7% and 87.9%. MIKGI also attained best performance in selecting VA local lab codes for desired laboratory tests and COVID-19 related features for COVID EHR studies. Compared to existing methods, MIKGI attained the most robust performance with accuracy the highest or near the highest across all tasks. CONCLUSIONS: The proposed MIKGI algorithm can effectively integrate incomplete summary data from biomedical text and EHR data to generate harmonized embeddings for EHR codes for knowledge graph modeling and cross-institutional translation of EHR codes. Doudou Zhou, Ziming Gan, Alina Patwari, Everett Neil Rush, Clara-Lea Bonzel, Vidul Ayakulangara Panickan, Chuan Hong, Yuk-Lam Ho, Tianrun A. Cai, Lauren Costa, Victor M. Castro, Shawn N. Murphy, Gabriel A. Brat, Griffin M. Weber, Paul Avillach, John Michael Gaziano, Kelly Cho, Katherine P. Liao, Tianxi Cai |
J. Biomed. Informatics | 14 |
| 2021 | Integrating Informatics for Integrating Biology and the Bedside with tranSMART: Flexible Data Warehousing with Complex Analytics
Jeffrey G. Klann, Michael Mendis, Peter Rice, Rudy Potenzone, Louisa May Klann, Griffin M. Weber, Diane Keogh, Shawn N. Murphy |
AMIA | 8 |
| 2021 | Combining Chart Review and Hospital System Dynamics for Electronic Health Record Phenotyping in an International COVID-19 Research Network
Jeffrey G. Klann, Griffin M. Weber, Emma Perez, William Yuan, Gabriel A. Brat, Shawn N. Murphy |
AMIA | 6 |
| 2021 | Temporal Phenotypic Pathways of Post-Acute Sequelae of SARS-CoV-2 by an International Consortium for Clinical Characterization of COVID-19 (4CE)
Shawn N. Murphy, Hossein Estiri, Arianna Dagliati, Riccardo Bellazzi, John H. Holmes |
AMIA | 1 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 7 |
| 2021 | A Machine Learning Approach for Identifying Emergent Phenotypes Associated with a Previous COVID Infection
Zachary H. Strasser, Hossein Estiri, Shawn N. Murphy |
AMIA | 3 |
| 2021 | A Novel Data Portal to Enable COVID-19 Data Integration and Analysis
Nich Wattanasin, Victor M. Castro, Vivian S. Gainer, Barbara Benoit, Andrew Cagan, Reeta Metta, Shawn N. Murphy |
AMIA | 7 |
| 2021 | High-throughput phenotyping with temporal sequencesabstractOBJECTIVE: High-throughput electronic phenotyping algorithms can accelerate translational research using data from electronic health record (EHR) systems. The temporal information buried in EHRs is often underutilized in developing computational phenotypic definitions. This study aims to develop a high-throughput phenotyping method, leveraging temporal sequential patterns from EHRs. MATERIALS AND METHODS: We develop a representation mining algorithm to extract 5 classes of representations from EHR diagnosis and medication records: the aggregated vector of the records (aggregated vector representation), the standard sequential patterns (sequential pattern mining), the transitive sequential patterns (transitive sequential pattern mining), and 2 hybrid classes. Using EHR data on 10 phenotypes from the Mass General Brigham Biobank, we train and validate phenotyping algorithms. RESULTS: Phenotyping with temporal sequences resulted in a superior classification performance across all 10 phenotypes compared with the standard representations in electronic phenotyping. The high-throughput algorithm's classification performance was superior or similar to the performance of previously published electronic phenotyping algorithms. We characterize and evaluate the top transitive sequences of diagnosis records paired with the records of risk factors, symptoms, complications, medications, or vaccinations. DISCUSSION: The proposed high-throughput phenotyping approach enables seamless discovery of sequential record combinations that may be difficult to assume from raw EHR data. Transitive sequences offer more accurate characterization of the phenotype, compared with its individual components, and reflect the actual lived experiences of the patients with that particular disease. CONCLUSION: Sequential data representations provide a precise mechanism for incorporating raw EHR records into downstream machine learning. Our approach starts with user interpretability and works backward to the technology. Hossein Estiri, Zachary H. Strasser, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Generative transfer learning for measuring plausibility of EHR diagnosis recordsabstractOBJECTIVE: Due to a complex set of processes involved with the recording of health information in the Electronic Health Records (EHRs), the truthfulness of EHR diagnosis records is questionable. We present a computational approach to estimate the probability that a single diagnosis record in the EHR reflects the true disease. MATERIALS AND METHODS: Using EHR data on 18 diseases from the Mass General Brigham (MGB) Biobank, we develop generative classifiers on a small set of disease-agnostic features from EHRs that aim to represent Patients, pRoviders, and their Interactions within the healthcare SysteM (PRISM features). RESULTS: We demonstrate that PRISM features and the generative PRISM classifiers are potent for estimating disease probabilities and exhibit generalizable and transferable distributional characteristics across diseases and patient populations. The joint probabilities we learn about diseases through the PRISM features via PRISM generative models are transferable and generalizable to multiple diseases. DISCUSSION: The Generative Transfer Learning (GTL) approach with PRISM classifiers enables the scalable validation of computable phenotypes in EHRs without the need for domain-specific knowledge about specific disease processes. CONCLUSION: Probabilities computed from the generative PRISM classifier can enhance and accelerate applied Machine Learning research and discoveries with EHR data. Hossein Estiri, Sebastien Vasey, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deploymentabstractOBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19. Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing |
J. Am. Medical Informatics Assoc. | 36 |
| 2021 | Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record dataabstractOBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites. Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 41 |
| 2021 | ATLAS: an automated association test using probabilistically linked health records with application to genetic studiesabstractOBJECTIVE: Large amounts of health data are becoming available for biomedical research. Synthesizing information across databases may capture more comprehensive pictures of patient health and enable novel research studies. When no gold standard mappings between patient records are available, researchers may probabilistically link records from separate databases and analyze the linked data. However, previous linked data inference methods are constrained to certain linkage settings and exhibit low power. Here, we present ATLAS, an automated, flexible, and robust association testing algorithm for probabilistically linked data. MATERIALS AND METHODS: Missing variables are imputed at various thresholds using a weighted average method that propagates uncertainty from probabilistic linkage. Next, estimated effect sizes are obtained using a generalized linear model. ATLAS then conducts the threshold combination test by optimally combining P values obtained from data imputed at varying thresholds using Fisher's method and perturbation resampling. RESULTS: In simulations, ATLAS controls for type I error and exhibits high power compared to previous methods. In a real-world genetic association study, meta-analysis of ATLAS-enabled analyses on a linked cohort with analyses using an existing cohort yielded additional significant associations between rheumatoid arthritis genetic risk score and laboratory biomarkers. DISCUSSION: Weighted average imputation weathers false matches and increases contribution of true matches to mitigate linkage error-induced bias. The threshold combination test avoids arbitrarily choosing a threshold to rule a match, thus automating linked data-enabled analyses and preserving power. CONCLUSION: ATLAS promises to enable novel and powerful research studies using linked data to capitalize on all available data sources. Harrison G. Zhang, Boris P. Hejblum, Griffin M. Weber, Nathan P. Palmer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Katherine P. Liao, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Multi-channel attention-fusion neural network for brain age estimation: Accuracy, generality, and interpretation with 16, 705 healthy MRIs across lifespan
Diana Pereira, Juan David Perez, Randy L. Gollub, Shawn N. Murphy, Sanjay Prabhu, Rudolph Pienaar, Richard Robertson, Patricia Ellen Grant, Yangming Ou |
Medical Image Anal. | 5 |
| 2020 | Transitive Sequential Pattern Mining for Discrete Clinical Data
Hossein Estiri, Sebastien Vasey, Shawn N. Murphy |
AIME | 3 |
| 2020 | Generative Transfer Learning for Measuring Plausibility of EHR Diagnosis Records Over Time
Hossein Estiri, Sebastien Vasey, Jeffrey G. Klann, Victor M. Castro, Shawn N. Murphy |
AMIA | 5 |
| 2020 | Informatics for Integrating Biology and the Bedside (i2b2) in 2020: Supporting Large Ontologies and REDcap Surveys
Jeffrey G. Klann, Michael Mendis, Diane Keogh, Shawn N. Murphy |
AMIA | 4 |
| 2020 | National Informatics Infrastructure Responding to the COVID-19 Pandemic
Douglas MacFadden, Shawn N. Murphy, Griffin M. Weber, Anupama Maram |
AMIA | 2 |
| 2020 | High-throughput Phenotyping with EHR Sequences
Shawn N. Murphy, Hossein Estiri, Zachary H. Strasser, Kavishwar B. Wagholikar, Victor M. Castro |
AMIA | 1 |
| 2020 | HistoriView: An Interactive and Scalable Visual Exploratory Plugin for Longitudinal Patient Data Review
Heekyong Park, Taowei David Wang, Vivian S. Gainer, Victor M. Castro, Nich Wattanasin, Shawn N. Murphy |
AMIA | 6 |
| 2020 | Polar labeling: silver standard algorithm for training disease classifiersabstractMOTIVATION: Expert-labeled data are essential to train phenotyping algorithms for cohort identification. However expert labeling is time and labor intensive, and the costs remain prohibitive for scaling phenotyping to wider use-cases. RESULTS: We present an approach referred to as polar labeling (PL), to create silver standard for training machine learning (ML) for disease classification. We test the hypothesis that ML models trained on the silver standard created by applying PL on unlabeled patient records, are comparable in performance to the ML models trained on gold standard, created by clinical experts through manual review of patient records. We perform experimental validation using health records of 38 023 patients spanning six diseases. Our results demonstrate the superior performance of the proposed approach. AVAILABILITY AND IMPLEMENTATION: We provide a Python implementation of the algorithm and the Python code developed for this study on Github. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kavishwar B. Wagholikar, Hossein Estiri, Marykate Murphy, Shawn N. Murphy |
Bioinform. | 4 |
| 2020 | sureLDA: A multidisease automated phenotyping method for the electronic health recordabstractOBJECTIVE: A major bottleneck hindering utilization of electronic health record data for translational research is the lack of precise phenotype labels. Chart review as well as rule-based and supervised phenotyping approaches require laborious expert input, hampering applicability to studies that require many phenotypes to be defined and labeled de novo. Though International Classification of Diseases codes are often used as surrogates for true labels in this setting, these sometimes suffer from poor specificity. We propose a fully automated topic modeling algorithm to simultaneously annotate multiple phenotypes. MATERIALS AND METHODS: Surrogate-guided ensemble latent Dirichlet allocation (sureLDA) is a label-free multidimensional phenotyping method. It first uses the PheNorm algorithm to initialize probabilities based on 2 surrogate features for each target phenotype, and then leverages these probabilities to constrain the LDA topic model to generate phenotype-specific topics. Finally, it combines phenotype-feature counts with surrogates via clustering ensemble to yield final phenotype probabilities. RESULTS: sureLDA achieves reliably high accuracy and precision across a range of simulated and real-world phenotypes. Its performance is robust to phenotype prevalence and relative informativeness of surogate vs nonsurrogate features. It also exhibits powerful feature selection properties. DISCUSSION: sureLDA combines attractive properties of PheNorm and LDA to achieve high accuracy and precision robust to diverse phenotype characteristics. It offers particular improvement for phenotypes insufficiently captured by a few surrogate features. Moreover, sureLDA's feature selection ability enables it to handle high feature dimensions and produce interpretable computational phenotypes. CONCLUSIONS: sureLDA is well suited toward large-scale electronic health record phenotyping for highly multiphenotype applications such as phenome-wide association studies . Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 7 |
| 2019 | sureLDA: A Novel Multi-Disease Automated Phenotyping Method for the Electronic Health Record
Yuri Ahuja, Doudou Zhou, Zeling He, Jiehuan Sun, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Chuan Hong, Tianxi Cai |
AMIA | 7 |
| 2019 | EHR Sequencing: A Novel Approach for Constructing Predictive and Interpretable Data Representations from EHR Data
Hossein Estiri, Thomas H. McCoy, Shawn N. Murphy |
AMIA | 3 |
| 2019 | Ontologies Enabling Computable Tables
Jeffrey G. Klann, Nich Wattanasin, Michael Mendis, Matthew A. Joss, Hossein Estiri, Kavishwar B. Wagholikar, Shawn N. Murphy |
AMIA | 7 |
| 2019 | Patient Stratification Process for enabling Clinical Interventions
Kavishwar B. Wagholikar, Samuel J. Aronson, Benjamin M. Scirica, Akshay S. Desai, Shawn N. Murphy |
AMIA | 5 |
| 2019 | Dynamic Phenotyping to facilitate Accrual for Prospective Clinical studies: A Case Study in Heart Failure
Kavishwar B. Wagholikar, Christina M. Fischer, Alyssa P. Goodson, Christopher Herrick, Taylor Maclean, Katelyn Smith, Liliana Fera, Thomas Gaziano, Jacqueline Dunning, Joshua Bosque-Hamilton, Lina Matta, Eloy Toscano, Brent Richter, Layne Ainsworth, Michael Oates, Samuel J. Aronson, Calum A. MacRae, Benjamin M. Scirica, Akshay S. Desai, Shawn N. Murphy |
AMIA | 20 |
| 2019 | Stratification of Patient Population for enabling data-driven Clinical Interventions using I2b2
Kavishwar B. Wagholikar, Vishal Vernekar, Akshay Zagade, Yuri Ostrovsky, Shek-Wayne Chan, Alyssa P. Goodson, Ameet Pathak, Corey Glynn, Christopher Herrick, Shawn N. Murphy |
AMIA | 10 |
| 2019 | Scalable Process to Generate Aggregated Patient Data for Analysis
Nich Wattanasin, Taowei David Wang, Vivian S. Gainer, Shawn N. Murphy |
AMIA | 4 |
| 2019 | Plugin for importing spreadsheets into Informatics for Integrating Biology and the Bedside platform
Akshay Zagade, Vishal Vernekar, Shek-Wayne Chan, Kavishwar B. Wagholikar, Rupendra Chulyadyo, Yuri Ostrovsky, Alyssa P. Goodson, Ameet Pathak, Christopher Herrick, Shawn N. Murphy |
AMIA | 10 |
| 2019 | A federated EHR network data completeness tracking systemabstractOBJECTIVE: The study sought to design, pilot, and evaluate a federated data completeness tracking system (CTX) for assessing completeness in research data extracted from electronic health record data across the Accessible Research Commons for Health (ARCH) Clinical Data Research Network. MATERIALS AND METHODS: The CTX applies a systems-based approach to design workflow and technology for assessing completeness across distributed electronic health record data repositories participating in a queryable, federated network. The CTX invokes 2 positive feedback loops that utilize open source tools (DQe-c and Vue) to integrate technology and human actors in a system geared for increasing capacity and taking action. A pilot implementation of the system involved 6 ARCH partner sites between January 2017 and May 2018. RESULTS: The ARCH CTX has enabled the network to monitor and, if needed, adjust its data management processes to maintain complete datasets for secondary use. The system allows the network and its partner sites to profile data completeness both at the network and partner site levels. Interactive visualizations presenting the current state of completeness in the context of the entire network as well as changes in completeness across time were valued among the CTX user base. DISCUSSION: Distributed clinical data networks are complex systems. Top-down approaches that solely rely on technology to report data completeness may be necessary but not sufficient for improving completeness (and quality) of data in large-scale clinical data networks. Improving and maintaining complete (high-quality) data in such complex environments entails sociotechnical systems that exploit technology and empower human actors to engage in the process of high-quality data curating. CONCLUSIONS: The CTX has increased the network's capacity to rapidly identify data completeness issues and empowered ARCH partner sites to get involved in improving the completeness of respective data in their repositories. Hossein Estiri, Jeffrey G. Klann, Sarah Weiler, Ernest Alema-Mensah, R. Joseph Applegate, Galina Lozinski, Nandan Patibandla, William G. Adams, Marc D. Natter, Elizabeth O. Ofili, Brian Ostasiewski, Alexander Quarshie, Gary E. Rosenthal, Elmer V. Bernstam, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 17 |
| 2019 | High-throughput multimodal automated phenotyping (MAP) with application to PheWASabstractOBJECTIVE: Electronic health records linked with biorepositories are a powerful platform for translational studies. A major bottleneck exists in the ability to phenotype patients accurately and efficiently. The objective of this study was to develop an automated high-throughput phenotyping method integrating International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP). MATERIALS AND METHODS: We developed a mapping method for automatically identifying relevant ICD and NLP concepts for a specific phenotype leveraging the Unified Medical Language System. Along with health care utilization, aggregated ICD and NLP counts were jointly analyzed by fitting an ensemble of latent mixture models. The multimodal automated phenotyping (MAP) algorithm yields a predicted probability of phenotype for each patient and a threshold for classifying participants with phenotype yes/no. The algorithm was validated using labeled data for 16 phenotypes from a biorepository and further tested in an independent cohort phenome-wide association studies (PheWAS) for 2 single nucleotide polymorphisms with known associations. RESULTS: The MAP algorithm achieved higher or similar AUC and F-scores compared to the ICD code across all 16 phenotypes. The features assembled via the automated approach had comparable accuracy to those assembled via manual curation (AUCMAP 0.943, AUCmanual 0.941). The PheWAS results suggest that the MAP approach detected previously validated associations with higher power when compared to the standard PheWAS method based on ICD codes. CONCLUSION: The MAP approach increased the accuracy of phenotype definition while maintaining scalability, thereby facilitating use in studies requiring large-scale phenotyping, such as PheWAS. Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Yuk-Lam Ho, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, Christopher J. O'Donnell, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002 |
J. Am. Medical Informatics Assoc. | 13 |
| 2019 | Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng |
J. Biomed. Informatics | 17 |
| 2018 | Using HL7 FHIR to Improve Standardization and Interoperability of Common Data Models for Clinical and Translational Research
Guoqian Jiang, Jon D. Duke, Daniella Meeker, Mitra Rocca, Harold R. Solbrig, Shawn N. Murphy |
AMIA | 6 |
| 2018 | Accessible Research Commons for Health: Four Years Into the PCORnet Journey
Jeffrey G. Klann, Stanley Boykin, Marc D. Natter, Margaret Vella, Douglas MacFadden, Sarah Weiler, Sebastian Schneeweiss, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 9 |
| 2018 | High-Throughput Multimodal Automated Phenotyping (MAP) Incorporating Natural Language Processing with Application to PheWAS
Katherine P. Liao, Jiehuan Sun, Tianrun A. Cai, Nicholas B. Link, Chuan Hong, Jie Huang 0030, Jennifer E. Huffman, Jessica L. Gronsbell, Lauren Costa, Victor M. Castro, Vivian S. Gainer, Shawn N. Murphy, John Michael Gaziano, Kelly Cho, Peter Szolovits, Isaac S. Kohane, Sheng Yu 0002, Tianxi Cai |
AMIA | 12 |
| 2018 | Automated Population of an i2b2 Clinical Data Warehouse using FHIR
Harold R. Solbrig, Na Hong, Shawn N. Murphy, Guoqian Jiang |
AMIA | 3 |
| 2018 | Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE exampleabstractThe eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians. Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor |
J. Am. Medical Informatics Assoc. | 14 |
| 2018 | Exploring completeness in clinical data research networks with DQe-cabstractObjective: To provide an open source, interoperable, and scalable data quality assessment tool for evaluation and visualization of completeness and conformance in electronic health record (EHR) data repositories. Materials and Methods: This article describes the tool's design and architecture and gives an overview of its outputs using a sample dataset of 200 000 randomly selected patient records with an encounter since January 1, 2010, extracted from the Research Patient Data Registry (RPDR) at Partners HealthCare. All the code and instructions to run the tool and interpret its results are provided in the Supplementary Appendix. Results: DQe-c produces a web-based report that summarizes data completeness and conformance in a given EHR data repository through descriptive graphics and tables. Results from running the tool on the sample RPDR data are organized into 4 sections: load and test details, completeness test, data model conformance test, and test of missingness in key clinical indicators. Discussion: Open science, interoperability across major clinical informatics platforms, and scalability to large databases are key design considerations for DQe-c. Iterative implementation of the tool across different institutions directed us to improve the scalability and interoperability of the tool and find ways to facilitate local setup. Conclusion: EHR data quality assessment has been hampered by implementation of ad hoc processes. The architecture and implementation of DQe-c offer valuable insights for developing reproducible and scalable data science tools to assess, manage, and process data in clinical data repositories. Hossein Estiri, Kari A. Stephens, Jeffrey G. Klann, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | Web services for data warehouses: OMOP and PCORnet on i2b2abstractObjective: Healthcare organizations use research data models supported by projects and tools that interest them, which often means organizations must support the same data in multiple models. The healthcare research ecosystem would benefit if tools and projects could be adopted independently from the underlying data model. Here, we introduce the concept of a reusable application programming interface (API) for healthcare and show that the i2b2 API can be adapted to support diverse patient-centric data models. Materials and Methods: We develop methodology for extending i2b2's pre-existing API to query additional data models, using i2b2's recent "multi-fact-table querying" feature. Our method involves developing data-model-specific i2b2 ontologies and mapping these to query non-standard table structure. Results: We implement this methodology to query OMOP and PCORnet models, which we validate with the i2b2 query tool. We implement the entire PCORnet data model and a five-domain subset of the OMOP model. We also demonstrate that additional, ancillary data model columns can be modeled and queried as i2b2 "modifiers." Discussion: i2b2's REST API can be used to query multiple healthcare data models, enabling shared tooling to have a choice of backend data stores. This enables separation between data model and software tooling for some of the more popular open analytic data models in healthcare. Conclusion: This methodology immediately allows querying OMOP and PCORnet using the i2b2 API. It is released as an open-source set of Docker images, and also on the i2b2 community wiki. Jeffrey G. Klann, Lori C. Phillips, Christopher Herrick, Matthew A. Joss, Kavishwar B. Wagholikar, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 6 |
| 2018 | Enabling phenotypic big data with PheNormabstractObjective: Electronic health record (EHR)-based phenotyping infers whether a patient has a disease based on the information in his or her EHR. A human-annotated training set with gold-standard disease status labels is usually required to build an algorithm for phenotyping based on a set of predictive features. The time intensiveness of annotation and feature curation severely limits the ability to achieve high-throughput phenotyping. While previous studies have successfully automated feature curation, annotation remains a major bottleneck. In this paper, we present PheNorm, a phenotyping algorithm that does not require expert-labeled samples for training. Methods: The most predictive features, such as the number of International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes or mentions of the target phenotype, are normalized to resemble a normal mixture distribution with high area under the receiver operating curve (AUC) for prediction. The transformed features are then denoised and combined into a score for accurate disease classification. Results: We validated the accuracy of PheNorm with 4 phenotypes: coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis. The AUCs of the PheNorm score reached 0.90, 0.94, 0.95, and 0.94 for the 4 phenotypes, respectively, which were comparable to the accuracy of supervised algorithms trained with sample sizes of 100-300, with no statistically significant difference. Conclusion: The accuracy of the PheNorm algorithms is on par with algorithms trained with annotated samples. PheNorm fully automates the generation of accurate phenotyping algorithms and demonstrates the capacity for EHR-driven annotations to scale to the next level - phenotypic big data. Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 9 |
| 2017 | Applying unsupervised learning to characterize rare observations in clinical data: the DQe-p tool
Hossein Estiri, Jeffrey G. Klann, Kavishwar B. Wagholikar, Shawn N. Murphy |
AMIA | 4 |
| 2017 | Integrating Patient Registries into Enterprise wide Patient-Discovery Strategies
Christopher Herrick, Alyssa P. Goodson, Wayne Chan, Lori C. Phillips, Michael Mendis, Shawn N. Murphy |
AMIA | 6 |
| 2017 | Reuse of PCORnet Data to Support the Precision Medicine Initiative: Data Model Harmonization
Jeffrey G. Klann, Matthew A. Joss, Kevin Embree, Shawn N. Murphy |
AMIA | 4 |
| 2017 | Web-Service-Enabled Apps for Research: SMART-on-FHIR for OMOP and PCORNet
Jeffrey G. Klann, Kavishwar B. Wagholikar, Lori C. Phillips, Matthew A. Joss, Shawn N. Murphy |
AMIA | 5 |
| 2017 | Building Better Timeline Interactions for Patient Chart Reviews
Heekyong Park, Taowei David Wang, Vivian S. Gainer, Victor M. Castro, Shawn N. Murphy |
AMIA | 5 |
| 2017 | High-throughput Phenotyping via Denoised Normal Mixture Transformation
Sheng Yu 0002, Yumeng Ma, Jessica L. Gronsbell, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
AMIA | 10 |
| 2017 | SMART-on-FHIR implemented over i2b2abstractWe have developed an interface to serve patient data from Informatics for Integrating Biology and the Bedside (i2b2) repositories in the Fast Healthcare Interoperability Resources (FHIR) format, referred to as a SMART-on-FHIR cell. The cell serves FHIR resources on a per-patient basis, and supports the "substitutable" modular third-party applications (SMART) OAuth2 specification for authorization of client applications. It is implemented as an i2b2 server plug-in, consisting of 6 modules: authentication, REST, i2b2-to-FHIR converter, resource enrichment, query engine, and cache. The source code is freely available as open source. We tested the cell by accessing resources from a test i2b2 installation, demonstrating that a SMART app can be launched from the cell that accesses patient data stored in i2b2. We successfully retrieved demographics, medications, labs, and diagnoses for test patients. The SMART-on-FHIR cell will enable i2b2 sites to provide simplified but secure data access in FHIR format, and will spur innovation and interoperability. Further, it transforms i2b2 into an apps platform. Kavishwar B. Wagholikar, Joshua C. Mandel, Jeffrey G. Klann, Nich Wattanasin, Michael Mendis, Christopher G. Chute, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 8 |
| 2017 | Biases introduced by filtering electronic health records for patients with "complete data"abstractOBJECTIVE: One promise of nationwide adoption of electronic health records (EHRs) is the availability of data for large-scale clinical research studies. However, because the same patient could be treated at multiple health care institutions, data from only a single site might not contain the complete medical history for that patient, meaning that critical events could be missing. In this study, we evaluate how simple heuristic checks for data "completeness" affect the number of patients in the resulting cohort and introduce potential biases. MATERIALS AND METHODS: We began with a set of 16 filters that check for the presence of demographics, laboratory tests, and other types of data, and then systematically applied all 216 possible combinations of these filters to the EHR data for 12 million patients at 7 health care systems and a separate payor claims database of 7 million members. RESULTS: EHR data showed considerable variability in data completeness across sites and high correlation between data types. For example, the fraction of patients with diagnoses increased from 35.0% in all patients to 90.9% in those with at least 1 medication. An unrelated claims dataset independently showed that most filters select members who are older and more likely female and can eliminate large portions of the population whose data are actually complete. DISCUSSION AND CONCLUSION: As investigators design studies, they need to balance their confidence in the completeness of the data with the effects of placing requirements on the data on the resulting patient cohort. Griffin M. Weber, William G. Adams, Elmer V. Bernstam, Jonathan P. Bickel, Kathe P. Fox, Keith Marsolo, Vijay A. Raghavan, Alexander Turchin, Shawn N. Murphy, Kenneth D. Mandl |
J. Am. Medical Informatics Assoc. | 10 |
| 2017 | Surrogate-assisted feature extraction for high-throughput phenotypingabstractOBJECTIVE: Phenotyping algorithms are capable of accurately identifying patients with specific phenotypes from within electronic medical records systems. However, developing phenotyping algorithms in a scalable way remains a challenge due to the extensive human resources required. This paper introduces a high-throughput unsupervised feature selection method, which improves the robustness and scalability of electronic medical record phenotyping without compromising its accuracy. METHODS: The proposed Surrogate-Assisted Feature Extraction (SAFE) method selects candidate features from a pool of comprehensive medical concepts found in publicly available knowledge sources. The target phenotype's International Classification of Diseases, Ninth Revision and natural language processing counts, acting as noisy surrogates to the gold-standard labels, are used to create silver-standard labels. Candidate features highly predictive of the silver-standard labels are selected as the final features. RESULTS: Algorithms were trained to identify patients with coronary artery disease, rheumatoid arthritis, Crohn's disease, and ulcerative colitis using various numbers of labels to compare the performance of features selected by SAFE, a previously published automated feature extraction for phenotyping procedure, and domain experts. The out-of-sample area under the receiver operating characteristic curve and F -score from SAFE algorithms were remarkably higher than those from the other two, especially at small label sizes. CONCLUSION: SAFE advances high-throughput phenotyping methods by automatically selecting a succinct set of informative features for algorithm training, which in turn reduces overfitting and the needed number of gold-standard labels. SAFE also potentially identifies important features missed by automated feature extraction for phenotyping or experts. Sheng Yu 0002, Abhishek Chakrabortty, Katherine P. Liao, Tianrun A. Cai, Ashwin N. Ananthakrishnan, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 9 |
| 2016 | Comparison of Data Models used in Research Data Repositories for Electronic Phenotyping
Jeffrey G. Klann, Vijay A. Raghavan, Michael Mendis, Douglas MacFadden, Sarah Weiler, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 7 |
| 2016 | Data Topography of a Large Multi-Site Research Network
Jeffrey G. Klann, Vijay A. Raghavan, Douglas MacFadden, Sarah Weiler, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 6 |
| 2016 | Cajun Codefest 4.0 on SMART-on-FHIR apps for Diabetes
Kavishwar B. Wagholikar, Eliel Oliveira, Henry Chu, Harshal Shah, Joshua C. Mandel, Jeffrey G. Klann, Sohail Rao, Kenneth D. Mandl, Shawn N. Murphy, Thomas Carton |
AMIA | 10 |
| 2016 | Evaluation of SMART-on-FHIR I2b2 cell using PCORNET data model
Kavishwar B. Wagholikar, Eliel Oliveira, Joshua C. Mandel, Jeffrey G. Klann, Prasad Patil, Kenneth D. Mandl, Shawn N. Murphy, Thomas Carton |
AMIA | 8 |
| 2016 | Export data from i2b2 using the new download data web client plugin
Nich Wattanasin, Taowei David Wang, Bhaswati Ghosh, Reeta Metta, Vivian S. Gainer, Shawn N. Murphy |
AMIA | 6 |
| 2016 | Data interchange using i2b2abstractOBJECTIVE: Reinventing data extraction from electronic health records (EHRs) to meet new analytical needs is slow and expensive. However, each new data research network that wishes to support its own analytics tends to develop its own data model. Joining these different networks without new data extraction, transform, and load (ETL) processes can reduce the time and expense needed to participate. The Informatics for Integrating Biology and the Bedside (i2b2) project supports data network interoperability through an ontology-driven approach. We use i2b2 as a hub, to rapidly reconfigure data to meet new analytical requirements without new ETL programming. MATERIALS AND METHODS: Our 12-site National Patient-Centered Clinical Research Network (PCORnet) Clinical Data Research Network (CDRN) uses i2b2 to query data. We developed a process to generate a PCORnet Common Data Model (CDM) physical database directly from existing i2b2 systems, thereby supporting PCORnet analytic queries without new ETL programming. This involved: a formalized process for representing i2b2 information models (the specification of data types and formats); an information model that represents CDM Version 1.0; and a program that generates CDM tables, driven by this information model. This approach is generalizable to any logical information model. RESULTS: Eight PCORnet CDRN sites have implemented this approach and generated a CDM database without a new ETL process from the EHR. This enables federated querying within the CDRN and compatibility with the national PCORnet Distributed Research Network. DISCUSSION: We have established a way to adapt i2b2 to new information models without requiring changes to the underlying data. Eight Scalable Collaborative Infrastructure for a Learning Health System sites vetted this methodology, resulting in a network that, at present, supports research on 10 million patients' data. CONCLUSION: New analytical requirements can be quickly and cost-effectively supported by i2b2 without creating new data extraction processes from the EHR. Jeffrey G. Klann, Aaron Abend, Vijay A. Raghavan, Kenneth D. Mandl, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 5 |
| 2015 | Demonstrating the Advantages of Applying Data Mining Techniques on Time-Dependent Electronic Medical Records
Uri Kartoun, Vishesh Kumar, Su-Chun Cheng, Sheng Yu 0002, Katherine P. Liao, Elizabeth W. Karlson, Ashwin N. Ananthakrishnan, Zongqi Xia, Vivian S. Gainer, Andrew Cagan, Guergana K. Savova, Pei J. Chen, Shawn N. Murphy, Susanne E. Churchill, Isaac S. Kohane, Peter Szolovits, Tianxi Cai, Stanley Y. Shaw |
AMIA | 13 |
| 2015 | The Scalable Collaborative Infrastructure for a Learning Health System
Jeffrey G. Klann, Marc D. Natter, Douglas MacFadden, Sarah Weiler, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 6 |
| 2015 | The Scalable Collaborative Infrastructure for a Learning Health System: Facilitating Agile Comparative Effectiveness Research
Jeffrey G. Klann, Marc D. Natter, Douglas MacFadden, Sarah Weiler, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 6 |
| 2015 | Supporting Multi-sourced Medication Information in i2b2
Jeffrey G. Klann, Pascal B. Pfiffner, Marc D. Natter, Emily Conner, Paul Blazejewski, Shawn N. Murphy, Kenneth D. Mandl |
AMIA | 6 |
| 2015 | SCILHS Data Mart Creation Plugin
Michael Mendis, Janice Donahoe, Jeffrey G. Klann, Vijay A. Raghavan, Lori C. Phillips, Alexander Turchin, Shawn N. Murphy |
AMIA | 7 |
| 2015 | Computable Phenotypes enabled by the i2b2 Validation Platform
Shawn N. Murphy, Vivian S. Gainer, Victor M. Castro, Alyssa P. Goodson, Lori C. Phillips, Sheng Yu 0002, Tianxi Cai |
AMIA | 1 |
| 2015 | (Authoring) Rules, (Distributed Query) Tools, and Drools: The challenging new world of high throughput phenotyping
Jennifer A. Pacheco, Abel N. Kho, Jyotishman Pathak, Joshua C. Denny, Shawn N. Murphy |
AMIA | 5 |
| 2015 | Taking advantage of continuity of care documents to populate a research repositoryabstractOBJECTIVE: Clinical data warehouses have accelerated clinical research, but even with available open source tools, there is a high barrier to entry due to the complexity of normalizing and importing data. The Office of the National Coordinator for Health Information Technology's Meaningful Use Incentive Program now requires that electronic health record systems produce standardized consolidated clinical document architecture (C-CDA) documents. Here, we leverage this data source to create a low volume standards based import pipeline for the Informatics for Integrating Biology and the Bedside (i2b2) clinical research platform. We validate this approach by creating a small repository at Partners Healthcare automatically from C-CDA documents. MATERIALS AND METHODS: We designed an i2b2 extension to import C-CDAs into i2b2. It is extensible to other sites with variances in C-CDA format without requiring custom code. We also designed new ontology structures for querying the imported data. RESULTS: We implemented our methodology at Partners Healthcare, where we developed an adapter to retrieve C-CDAs from Enterprise Services. Our current implementation supports demographics, encounters, problems, and medications. We imported approximately 17 000 clinical observations on 145 patients into i2b2 in about 24 min. We were able to perform i2b2 cohort finding queries and view patient information through SMART apps on the imported data. DISCUSSION: This low volume import approach can serve small practices with local access to C-CDAs and will allow patient registries to import patient supplied C-CDAs. These components will soon be available open source on the i2b2 wiki. CONCLUSIONS: Our approach will lower barriers to entry in implementing i2b2 where informatics expertise or data access are limited. Jeffrey G. Klann, Michael Mendis, Lori C. Phillips, Alyssa P. Goodson, Beatriz H. S. C. Rocha, Howard Goldberg, Nich Wattanasin, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 8 |
| 2015 | Toward high-throughput phenotyping: unbiased automated feature extraction and selection from knowledge sourcesabstractOBJECTIVE: Analysis of narrative (text) data from electronic health records (EHRs) can improve population-scale phenotyping for clinical and genetic research. Currently, selection of text features for phenotyping algorithms is slow and laborious, requiring extensive and iterative involvement by domain experts. This paper introduces a method to develop phenotyping algorithms in an unbiased manner by automatically extracting and selecting informative features, which can be comparable to expert-curated ones in classification accuracy. MATERIALS AND METHODS: Comprehensive medical concepts were collected from publicly available knowledge sources in an automated, unbiased fashion. Natural language processing (NLP) revealed the occurrence patterns of these concepts in EHR narrative notes, which enabled selection of informative features for phenotype classification. When combined with additional codified features, a penalized logistic regression model was trained to classify the target phenotype. RESULTS: The authors applied our method to develop algorithms to identify patients with rheumatoid arthritis and coronary artery disease cases among those with rheumatoid arthritis from a large multi-institutional EHR. The area under the receiver operating characteristic curves (AUC) for classifying RA and CAD using models trained with automated features were 0.951 and 0.929, respectively, compared to the AUCs of 0.938 and 0.929 by models trained with expert-curated features. DISCUSSION: Models trained with NLP text features selected through an unbiased, automated procedure achieved comparable or slightly higher accuracy than those trained with expert-curated features. The majority of the selected model features were interpretable. CONCLUSION: The proposed automated feature extraction method, generating highly accurate phenotyping algorithms with improved efficiency, is a significant step toward high-throughput phenotyping. Sheng Yu 0002, Katherine P. Liao, Stanley Y. Shaw, Vivian S. Gainer, Susanne E. Churchill, Peter Szolovits, Shawn N. Murphy, Isaac S. Kohane, Tianxi Cai |
J. Am. Medical Informatics Assoc. | 7 |
| 2014 | Integrating Information from Unstructured Text with Structured Clinical Data from an Electronic Medical Record to Improve Patient Cohort Identification
Victor M. Castro, Sergey Goryachev, Christopher Herrick, Vivian S. Gainer, Martin Rees, Shawn N. Murphy |
AMIA | 6 |
| 2014 | Enabling Patient-Centric Comparative Effectiveness Research in i2b2
Jeffrey G. Klann, Lori C. Phillips, Kenneth D. Mandl, Shawn N. Murphy |
AMIA | 4 |
| 2014 | Handling Clinical and Next Generation Sequencing data: new strategies in i2b2 and tranSMART
Shawn N. Murphy, Riccardo Bellazzi, Matteo Gabetta, Paul Avillach, Lori C. Phillips |
AMIA | 1 |
| 2014 | Informatics for Integrating Biology and the Bedside (I2b2) Clinical Trials (CT) Patient Ascertainment Suite
Shawn N. Murphy, Nich Wattanasin, Susanne E. Churchill, Isaac S. Kohane, Vivian S. Gainer |
AMIA | 1 |
| 2014 | Improving the Review of Individual Patients in a Clinical Data Repository
Nich Wattanasin, Michael Mendis, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy |
AMIA | 5 |
| 2014 | Query Health: standards-based, cross-platform population health surveillanceabstractOBJECTIVE: Understanding population-level health trends is essential to effectively monitor and improve public health. The Office of the National Coordinator for Health Information Technology (ONC) Query Health initiative is a collaboration to develop a national architecture for distributed, population-level health queries across diverse clinical systems with disparate data models. Here we review Query Health activities, including a standards-based methodology, an open-source reference implementation, and three pilot projects. MATERIALS AND METHODS: Query Health defined a standards-based approach for distributed population health queries, using an ontology based on the Quality Data Model and Consolidated Clinical Document Architecture, Health Quality Measures Format (HQMF) as the query language, the Query Envelope as the secure transport layer, and the Quality Reporting Document Architecture as the result language. RESULTS: We implemented this approach using Informatics for Integrating Biology and the Bedside (i2b2) and hQuery for data analytics and PopMedNet for access control, secure query distribution, and response. We deployed the reference implementation at three pilot sites: two public health departments (New York City and Massachusetts) and one pilot designed to support Food and Drug Administration post-market safety surveillance activities. The pilots were successful, although improved cross-platform data normalization is needed. DISCUSSIONS: This initiative resulted in a standards-based methodology for population health queries, a reference implementation, and revision of the HQMF standard. It also informed future directions regarding interoperability and data access for ONC's Data Access Framework initiative. CONCLUSIONS: Query Health was a test of the learning health system that supplied a functional methodology and reference implementation for distributed population health queries that has been validated at three sites. Jeffrey G. Klann, Michael D. Buck, Jeffrey S. Brown, Marc Hadley, Richard Elmore, Griffin M. Weber, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 7 |
| 2014 | Brief communication: Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS): ArchitectureabstractWe describe the architecture of the Patient Centered Outcomes Research Institute (PCORI) funded Scalable Collaborative Infrastructure for a Learning Healthcare System (SCILHS, http://www.SCILHS.org) clinical data research network, which leverages the $48 billion dollar federal investment in health information technology (IT) to enable a queryable semantic data model across 10 health systems covering more than 8 million patients, plugging universally into the point of care, generating evidence and discovery, and thereby enabling clinician and patient participation in research during the patient encounter. Central to the success of SCILHS is development of innovative 'apps' to improve PCOR research methods and capacitate point of care functions such as consent, enrollment, randomization, and outreach for patient-reported outcomes. SCILHS adapts and extends an existing national research network formed on an advanced IT infrastructure built with open source, free, modular components. Kenneth D. Mandl, Isaac S. Kohane, Douglas MacFadden, Griffin M. Weber, Marc D. Natter, Joshua C. Mandel, Sebastian Schneeweiss, Sarah Weiler, Jeffrey G. Klann, Jonathan P. Bickel, William G. Adams, Yaorong Ge, James Perkins, Keith Marsolo, Elmer V. Bernstam, John Showalter, Alexander Quarshie, Elizabeth O. Ofili, George Hripcsak, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 21 |
| 2014 | Evaluation of matched control algorithms in EHR-based phenotyping studies: A case study of inflammatory bowel disease comorbidities
Victor M. Castro, W. Kay Apperson, Vivian S. Gainer, Ashwin N. Ananthakrishnan, Alyssa P. Goodson, Taowei David Wang, Christopher Herrick, Shawn N. Murphy |
J. Biomed. Informatics | 8 |
| 2013 | Query Health: One Step Toward a Learning Health System
Jeffrey G. Klann, Michael D. Buck, Jeffrey S. Brown, Shawn N. Murphy, Douglas B. Fridsma |
AMIA | 4 |
| 2013 | Using SMART and i2b2 to Efficiently Identify Adverse Events
Jeffrey G. Klann, Rachel Badovinac Ramoni, Shawn N. Murphy |
AMIA | 3 |
| 2013 | Amassing Pediatric Brain MRI's to Understand "Normal" using Mi2b2
Shawn N. Murphy, Christopher Herrick, Victor M. Castro, Randy L. Gollub, Nathaniel Reynolds, Patricia Ellen Grant |
AMIA | 1 |
| 2013 | Organization and Transformation of Next-Generation Sequencing Data for Use within i2b2
Lori C. Phillips, Shawn N. Murphy, Isaac S. Kohane |
AMIA | 2 |
| 2013 | Integrating the CCDA for Real-Time Patient Data in the i2b2 Platform
Nich Wattanasin, Michael Mendis, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy |
AMIA | 7 |
| 2013 | Research Informatics : Re-engineering the Research Enterprise
Mark G. Weiner, Philip R. O. Payne, Peter J. Embí, Shawn N. Murphy |
AMIA | 4 |
| 2012 | Data Sharing: Incentives and Governance Issues in Industry and Academia
Suzanne Bakken, Michael N. Cantor, Shawn N. Murphy, Lisa M. Schilling |
AMIA | 3 |
| 2012 | Implementing a pharmacovigilance framework using data from electronic medical records
Victor M. Castro, Vivian S. Gainer, Christopher Herrick, Shawn N. Murphy, Wannapa Kay Mahamaneerat |
AMIA | 4 |
| 2012 | Query Health and i2b2: Enabling Standards-based, Multiplatform Population Health Queries
Jeffrey G. Klann, Shawn N. Murphy |
AMIA | 2 |
| 2012 | The Medical App Store, Research Data Repositories, and Physician Cognitive Overload: Uniting Three Large, Multisite Grants for Health Care Transformation
Jeffrey G. Klann, Adam Wright, Allison B. McCoy, Shawn N. Murphy |
AMIA | 4 |
| 2012 | Building the SMART Platforms Ecosystem: Toward an Apps-Based Health Information Economy
Kenneth D. Mandl, Brian D. Athey, Daniel S. Fritsch, Shawn N. Murphy, Will Ross |
AMIA | 4 |
| 2012 | Supporting Population Queries and Clinical Trials in i2b2 with SMART
Shawn N. Murphy, Michael Mendis, Nich Wattanasin, Alyssa Porter, Stella Ubaha, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane |
AMIA | 1 |
| 2012 | Rethinking the "Honest Broker" in the Changing Face of Security and Privacy
Luke V. Rasmussen, Brian D. Athey, Andrew D. Boyd, Bradley A. Malin, Shawn N. Murphy |
AMIA | 5 |
| 2012 | Apps to display patient data, making SMART available in the i2b2 platform
Nich Wattanasin, Alyssa Porter, Stella Ubaha, Michael Mendis, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy |
AMIA | 10 |
| 2012 | Integrating Substitutable Medical Apps, Reusable Technologies (SMART) in the i2b2 Platform
Nich Wattanasin, Alyssa Porter, Stella Ubaha, Michael Mendis, Lori C. Phillips, Joshua C. Mandel, Rachel Badovinac Ramoni, Kenneth D. Mandl, Isaac S. Kohane, Shawn N. Murphy |
AMIA | 10 |
| 2012 | Clinical Bioinformatics: challenges and opportunitiesabstractBACKGROUND: Network Tools and Applications in Biology (NETTAB) Workshops are a series of meetings focused on the most promising and innovative ICT tools and to their usefulness in Bioinformatics. The NETTAB 2011 workshop, held in Pavia, Italy, in October 2011 was aimed at presenting some of the most relevant methods, tools and infrastructures that are nowadays available for Clinical Bioinformatics (CBI), the research field that deals with clinical applications of bioinformatics. METHODS: In this editorial, the viewpoints and opinions of three world CBI leaders, who have been invited to participate in a panel discussion of the NETTAB workshop on the next challenges and future opportunities of this field, are reported. These include the development of data warehouses and ICT infrastructures for data sharing, the definition of standards for sharing phenotypic data and the implementation of novel tools to implement efficient search computing solutions. RESULTS: Some of the most important design features of a CBI-ICT infrastructure are presented, including data warehousing, modularity and flexibility, open-source development, semantic interoperability, integrated search and retrieval of -omics information. CONCLUSIONS: Clinical Bioinformatics goals are ambitious. Many factors, including the availability of high-throughput "-omics" technologies and equipment, the widespread availability of clinical data warehouses and the noteworthy increase in data storage and computational power of the most recent ICT systems, justify research and efforts in this domain, which promises to be a crucial leveraging factor for biomedical research. Riccardo Bellazzi, Marco Masseroli, Shawn N. Murphy, Amnon Shabo, Paolo Romano 0001 |
BMC Bioinform. | 3 |
| 2012 | Portability of an algorithm to identify rheumatoid arthritis in electronic health recordsabstractOBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining. Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 14 |
| 2012 | A translational engine at the national scale: informatics for integrating biology and the bedsideabstractInformatics for integrating biology and the bedside (i2b2) seeks to provide the instrumentation for using the informational by-products of health care and the biological materials accumulated through the delivery of health care to conduct discovery research and to study the healthcare system in vivo. This complements existing efforts such as prospective cohort studies or trials outside the delivery of routine health care. i2b2 has been used to generate genome-wide studies at less than one tenth the cost and one tenth the time of conventionally performed studies as well as to identify important risk from commonly used medications. i2b2 has been adopted by over 60 academic health centers internationally. Isaac S. Kohane, Susanne E. Churchill, Shawn N. Murphy |
J. Am. Medical Informatics Assoc. | 3 |
| 2012 | The SMART Platform: early experience enabling substitutable applications for electronic health recordsabstractOBJECTIVE: The Substitutable Medical Applications, Reusable Technologies (SMART) Platforms project seeks to develop a health information technology platform with substitutable applications (apps) constructed around core services. The authors believe this is a promising approach to driving down healthcare costs, supporting standards evolution, accommodating differences in care workflow, fostering competition in the market, and accelerating innovation. MATERIALS AND METHODS: The Office of the National Coordinator for Health Information Technology, through the Strategic Health IT Advanced Research Projects (SHARP) Program, funds the project. The SMART team has focused on enabling the property of substitutability through an app programming interface leveraging web standards, presenting predictable data payloads, and abstracting away many details of enterprise health information technology systems. Containers--health information technology systems, such as electronic health records (EHR), personally controlled health records, and health information exchanges that use the SMART app programming interface or a portion of it--marshal data sources and present data simply, reliably, and consistently to apps. RESULTS: The SMART team has completed the first phase of the project (a) defining an app programming interface, (b) developing containers, and (c) producing a set of charter apps that showcase the system capabilities. A focal point of this phase was the SMART Apps Challenge, publicized by the White House, using http://www.challenge.gov website, and generating 15 app submissions with diverse functionality. CONCLUSION: Key strategic decisions must be made about the most effective market for further disseminating SMART: existing market-leading EHR vendors, new entrants into the EHR market, or other stakeholders such as health information exchanges. Kenneth D. Mandl, Joshua C. Mandel, Shawn N. Murphy, Elmer V. Bernstam, Rachel Badovinac Ramoni, David A. Kreda, J. Michael McCoy, Ben Adida, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | Strategies for maintaining patient privacy in i2b2abstractBACKGROUND: The re-use of patient data from electronic healthcare record systems can provide tremendous benefits for clinical research, but measures to protect patient privacy while utilizing these records have many challenges. Some of these challenges arise from a misperception that the problem should be solved technically when actually the problem needs a holistic solution. OBJECTIVE: The authors' experience with informatics for integrating biology and the bedside (i2b2) use cases indicates that the privacy of the patient should be considered on three fronts: technical de-identification of the data, trust in the researcher and the research, and the security of the underlying technical platforms. METHODS: The security structure of i2b2 is implemented based on consideration of all three fronts. It has been supported with several use cases across the USA, resulting in five privacy categories of users that serve to protect the data while supporting the use cases. RESULTS: The i2b2 architecture is designed to provide consistency and faithfully implement these user privacy categories. These privacy categories help reflect the policy of both the Health Insurance Portability and Accountability Act and the provisions of the National Research Act of 1974, as embodied by current institutional review boards. CONCLUSION: By implementing a holistic approach to patient privacy solutions, i2b2 is able to help close the gap between principle and practice. Shawn N. Murphy, Vivian S. Gainer, Michael Mendis, Susanne E. Churchill, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2)abstractInformatics for Integrating Biology and the Bedside (i2b2) is one of seven projects sponsored by the NIH Roadmap National Centers for Biomedical Computing (http://www.ncbcs.org). Its mission is to provide clinical investigators with the tools necessary to integrate medical record and clinical research data in the genomics age, a software suite to construct and integrate the modern clinical research chart. i2b2 software may be used by an enterprise's research community to find sets of interesting patients from electronic patient medical record data, while preserving patient privacy through a query tool interface. Project-specific mini-databases ("data marts") can be created from these sets to make highly detailed data available on these specific patients to the investigators on the i2b2 platform, as reviewed and restricted by the Institutional Review Board. The current version of this software has been released into the public domain and is available at the URL: http://www.i2b2.org/software. Shawn N. Murphy, Griffin M. Weber, Michael Mendis, Vivian S. Gainer, Henry C. Chueh, Susanne E. Churchill, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 1 |
| 2009 | Application of Information Technology: The Shared Health Research Information Network (SHRINE): A Prototype Federated Query Tool for Clinical Data RepositoriesabstractThe authors developed a prototype Shared Health Research Information Network (SHRINE) to identify the technical, regulatory, and political challenges of creating a federated query tool for clinical data repositories. Separate Institutional Review Boards (IRBs) at Harvard's three largest affiliated health centers approved use of their data, and the Harvard Medical School IRB approved building a Query Aggregator Interface that can simultaneously send queries to each hospital and display aggregate counts of the number of matching patients. Our experience creating three local repositories using the open source Informatics for Integrating Biology and the Bedside (i2b2) platform can be used as a road map for other institutions. The authors are actively working with the IRBs and regulatory groups to develop procedures that will ultimately allow investigators to obtain identified patient data and biomaterials through SHRINE. This will guide us in creating a future technical architecture that is scalable to a national level, compliant with ethical guidelines, and protective of the interests of the participating hospitals. Griffin M. Weber, Shawn N. Murphy, Andrew J. McMurry, Douglas MacFadden, Daniel J. Nigrin, Susanne E. Churchill, Isaac S. Kohane |
J. Am. Medical Informatics Assoc. | 2 |
| 2008 | Aligning temporal data by sentinel events: discovering patterns in electronic health recordsabstractElectronic Health Records (EHRs) and other temporal databases contain hidden patterns that reveal important cause-and-effect phenomena. Finding these patterns is a challenge when using traditional query languages and tabular displays. We present an interactive visual tool that complements query formulation by providing operations to align, rank and filter the results, and to visualize estimates of the intervals of validity of the data. Display of patient histories aligned on sentinel events (such as a first heart attack) enables users to spot precursor, co-occurring, and aftereffect events. A controlled study demonstrates the benefits of providing alignment (with a 61% speed improvement for complex tasks). A qualitative study and interviews with medical professionals demonstrates that the interface can be learned quickly and seems to address their needs. Taowei David Wang, Catherine Plaisant, Alexander J. Quinn, Roman Stanchak, Shawn N. Murphy, Ben Shneiderman |
CHI | 5 |
| 2008 | A National Human Neuroimaging Collaboratory Enabled by the Biomedical Informatics Research Network (BIRN)abstractThe aggregation of imaging, clinical, and behavioral data from multiple independent institutions and researchers presents both a great opportunity for biomedical research as well as a formidable challenge. Many research groups have well-established data collection and analysis procedures, as well as data and metadata format requirements that are particular to that group. Moreover, the types of data and metadata collected are quite diverse, including image, physiological, and behavioral data, as well as descriptions of experimental design, and preprocessing and analysis methods. Each of these types of data utilizes a variety of software tools for collection, storage, and processing. Furthermore sites are reluctant to release control over the distribution and access to the data and the tools. To address these needs, the Biomedical Informatics Research Network (BIRN) has developed a federated and distributed infrastructure for the storage, retrieval, analysis, and documentation of biomedical imaging data. The infrastructure consists of distributed data collections hosted on dedicated storage and computational resources located at each participating site, a federated data management system and data integration environment, an Extensible Markup Language (XML) schema for data exchange, and analysis pipelines, designed to leverage both the distributed data management environment and the available grid computing resources. David B. Keator, Jeffrey S. Grethe, Daniel S. Marcus, Ibrahim Burak Özyurt, Syam Gadde, Shawn N. Murphy, Steven D. Pieper, Douglas N. Greve, R. Notestine, Henry Jeremy Bockholt, Philip M. Papadopoulos |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2007 | Architecture of the Open-source Clinical Research Chart from Informatics for Integrating Biology and the Bedside
Shawn N. Murphy, Michael Mendis, Kristel Hackett, Rajesh Kuttan, Wensong Pan, Lori C. Phillips, Vivian S. Gainer, David Berkowicz, John P. Glaser, Isaac S. Kohane, Henry C. Chueh |
AMIA | 1 |
| 2006 | Mapping Tool for Maintaining Vocabulary Relationships
Kristel Hackett, Isaac S. Kohane, Henry C. Chueh, Shawn N. Murphy |
AMIA | 4 |
| 2006 | Integration of Clinical and Genetic Data in the i2b2 Architecture
Shawn N. Murphy, Michael Mendis, David A. Berkowitz, Isaac S. Kohane, Henry C. Chueh |
AMIA | 1 |
| 2006 | A Web Portal that Enables Collaborative Use of Advanced Medical Image Processing and Informatics Tools through the Biomedical Informatics Research Network (BIRN)
Shawn N. Murphy, Michael Mendis, Jeffrey S. Grethe, Randy L. Gollub, David N. Kennedy, Bruce R. Rosen |
AMIA | 1 |
| 2006 | Calculating the Benefits of a Research Patient Data Repository
Ruth Nalichowski, Diane Keogh, Henry C. Chueh, Shawn N. Murphy |
AMIA | 4 |
| 2005 | Concept-Value Pair Extraction from Semi-Structured Clinical Narrative: A Case Study Using Echocardiogram Reports
Jeanhee Chung, Shawn N. Murphy |
AMIA | 2 |
| 2003 | A Visual Interface Designed for Novice Users to find Research Patient Cohorts in a Large Biomedical Database
Shawn N. Murphy, Vivian S. Gainer, Henry C. Chueh |
AMIA | 1 |
| 2002 | A security architecture for query tools used to access large biomedical databases
Shawn N. Murphy, Henry C. Chueh |
AMIA | 1 |
| 2001 | Visual Query Tool for Integrating Clinical and Genetic Data in the Partners Healthcare System
Shawn N. Murphy, Henry C. Chueh |
AMIA | 1 |
| 2000 | Visual query tool for finding patient cohorts from a clinical data warehouse of the partners HealthCare system
Shawn N. Murphy, G. Octo Barnett, Henry C. Chueh |
AMIA | 1 |
| 1999 | Deriving Patient Profiles from Clinical Repositories
Rajneesh Behal, Shawn N. Murphy, G. Octo Barnett, Henry C. Chueh |
AMIA | 2 |
| 1999 | Optimizing healthcare research data warehouse design through past COSTAR query analysis
Shawn N. Murphy, Mary Morgan, G. Octo Barnett, Henry C. Chueh |
AMIA | 1 |
| 1998 | Using web technology and Java mobile software agents to manage outside referrals
Shawn N. Murphy, T. Ng, Dean F. Sittig, G. Octo Barnett |
AMIA | 1 |
| 1998 | Research Paper: The GuideLine Interchange Format: A Model for Representing GuidelinesabstractOBJECTIVE: To allow exchange of clinical practice guidelines among institutions and computer-based applications. DESIGN: The GuideLine Interchange Format (GLIF) specification consists of GLIF model and the GLIF syntax. The GLIF model is an object-oriented representation that consists of a set of classes for guideline entities, attributes for those classes, and data types for the attribute values. The GLIF syntax specifies the format of the test file that contains the encoding. METHODS: Researchers from the InterMed Collaboratory at Columbia University, Harvard University (Brigham and Women's Hospital and Massachusetts General Hospital), and Stanford University analyzed four existing guideline systems to derive a set of requirements for guideline representation. The GLIF specification is a consensus representation developed through a brainstorming process. Four clinical guidelines were encoded in GLIF to assess its expressivity and to study the variability that occurs when two people from different sites encode the same guideline. RESULTS: The encoders reported that GLIF was adequately expressive. A comparison of the encodings revealed substantial variability. CONCLUSION: GLIF was sufficient to model the guidelines for the four conditions that were examined. GLIF needs improvement in standard representation of medical concepts, criterion logic, temporal information, and uncertainty. Lucila Ohno-Machado, John H. Gennari, Shawn N. Murphy, Nilesh L. Jain, Samson W. Tu, Diane E. Oliver, Edward Pattison-Gordon, Robert A. Greenes, Edward H. Shortliffe, G. Octo Barnett |
J. Am. Medical Informatics Assoc. | 3 |
| 1997 | Puya: a method of attracting attention to relevant physical findings
W. D. de Estrada, Shawn N. Murphy, G. Octo Barnett |
AMIA | 2 |
| 1997 | Using software agents to maintain autonomous patient registries for clinical research
Shawn N. Murphy, Usman Rabbani, G. Octo Barnett |
AMIA | 1 |