Nigam H. Shah

dblp:s/NHShah · also N. H. Shah 0001, Nigam Shah · DBLP profile ↗
← Back
103ranked-venue papers
12as first author
32since 2021 · last 2025
0000-0001-9385-7158ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 83 · 10 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Time-to-Event Pretraining for 3D Medical Imaging
abstract
With the rise of medical foundation models and the growing availability of imaging data, scalable pretraining techniques offer a promising way to identify imaging biomarkers predictive of future disease risk. While current self-supervised methods for 3D medical imaging models capture local structural features like organ morphology, they fail to link pixel biomarkers with long-term health outcomes due to a missing context problem. Current approaches lack the temporal context necessary to identify biomarkers correlated with disease progression, as they rely on supervision derived only from images and concurrent text descriptions. To address this, we introduce time-to-event pretraining, a pretraining framework for 3D medical imaging models that leverages large-scale temporal supervision from paired, longitudinal electronic health records (EHRs). Using a dataset of 18,945 CT scans (4.2 million 2D images) and time-to-event distributions across thousands of EHR-derived tasks, our method improves outcome prediction, achieving an average AUROC increase of 23.7% and a 29.4% gain in Harrell’s C-index across 8 benchmark tasks. Importantly, these gains are achieved without sacrificing diagnostic classification performance. This study lays the foundation for integrating longitudinal EHR and 3D imaging data to advance clinical risk prediction.
Zepeng Huo, Jason Alan Fries, Alejandro Lozano, Jeya Maria Jose Valanarasu, Ethan Steinberg, Louis Blankemeier, Akshay Chaudhari, Curt Langlotz, Nigam H. Shah
ICLR9
2025 Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHR Data
abstract
Foundation Models (FMs) trained on Electronic Health Records (EHRs) have achieved state-of-the-art results on numerous clinical prediction tasks. However, prior EHR FMs typically have context windows of $<$1k tokens, which prevents them from modeling full patient EHRs which can exceed 10k's of events. For making clinical predictions, both model performance and robustness to the unique properties of EHR data are crucial. Recent advancements in subquadratic long-context architectures (e.g. Mamba) offer a promising solution. However, their application to EHR data has not been well-studied. We address this gap by presenting the first systematic evaluation of the effect of context length on modeling EHR data. We find that longer context models improve predictive performance -- our Mamba-based model surpasses the prior state-of-the-art on 9/14 tasks on the EHRSHOT prediction benchmark. Additionally, we measure robustness to three unique, previously underexplored properties of EHR data: (1) the prevalence of ``copy-forwarded" diagnoses which create artificial token repetition in EHR sequences; (2) the irregular time intervals between EHR events which can lead to a wide range of timespans within a context window; and (3) the natural increase in disease complexity over time which makes later tokens in the EHR harder to predict than earlier ones. Stratifying our EHRSHOT results, we find that higher levels of each property correlate negatively with model performance (e.g., a 14% higher Brier loss between the least and most irregular patients), but that longer context models are more robust to more extreme levels of these properties. Our work highlights the potential for using long-context architectures to model EHR data, and offers a case study on how to identify and quantify new challenges in modeling sequential data motivated by domains outside of natural language. We release all of our model checkpoints and code.
Michael Wornow, Suhana Bedi, Miguel Angel Fuentes Hernandez, Ethan Steinberg, Jason Alan Fries, Christopher Ré, Oluwasanmi Koyejo, Nigam H. Shah
ICLR8
2025 STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
abstract
Multi-class tissue-type classification of colorectal cancer (CRC) histopathologic images is a significant step in the development of downstream machine learning models for diagnosis and treatment planning. However, publicly available CRC datasets used to build tissue classifiers often suffer from insufficient morphologic diversity, class imbalance, and low-quality image tiles, limiting downstream model performance and generalizability. To address this research gap, we introduce STARC-9 (STAnford coloRectal Cancer), a large-scale dataset for multi-class tissue classification. STARC-9 comprises 630,000 histopathologic image tiles uniformly sampled across nine clinically relevant tissue classes (each represented by 70,000 tiles), systematically extracted from hematoxylin & eosin-stained whole-slide images (WSI) from 200 CRC patients at the Stanford University School of Medicine. To construct STARC-9, we propose a novel framework, DeepCluster++, consisting of two primary steps to ensure diversity within each tissue class, followed by pathologist verification. First, an encoder from an autoencoder trained specifically on histopathologic images is used to extract feature vectors from all tiles within a given input WSI. Next, K-means clustering groups morphologically similar tiles, followed by an equal-frequency binning method to sample diverse patterns within each tissue class. Finally, the selected tiles are verified by expert gastrointestinal pathologists to ensure classification accuracy. This semi-automated approach significantly reduces the manual effort required for dataset curation while producing high-quality training examples. To validate the utility of STARC-9, we benchmarked baseline convolutional neural networks, transformers, and pathology-specific foundation models on downstream multi-class CRC tissue classification and segmentation tasks when trained on STARC-9 versus publicly available datasets, demonstrating superior generalizability of models trained on STARC-9. Although we demonstrate the utility of DeepCluster++ on CRC as a pilot use-case, it is a flexible framework that can be used for constructing high-quality datasets from large WSI repositories across a wide range of cancer and non-cancer applications.
Barathi Subramanian, Rathinaraja Jeyaraj, Mitchell Nevin Peterson, Terry Guo, Nigam H. Shah, Curt Langlotz, Andrew Y. Ng, Jeanne Shen
NeurIPS5
2025 Reformulating patient stratification for targeting interventions by accounting for severity of downstream outcomes resulting from disease onset: a case study in sepsis
abstract
OBJECTIVES: To quantify differences between (1) stratifying patients by predicted disease onset risk alone and (2) stratifying by predicted disease onset risk and severity of downstream outcomes. We perform a case study of predicting sepsis. MATERIALS AND METHODS: We performed a retrospective analysis using observational data from Michigan Medicine at the University of Michigan (U-M) between 2016 and 2020 and the Beth Israel Deaconess Medical Center (BIDMC) between 2008 and 2012. We measured the correlation between the estimated sepsis risk and the estimated effect of sepsis on mortality using Spearman's correlation. We compared patients stratified by sepsis risk with patients stratified by sepsis risk and effect of sepsis on mortality. RESULTS: The U-M and BIDMC cohorts included 7282 and 5942 ICU visits; 7.9% and 8.1% developed sepsis, respectively. Among visits with sepsis, 21.9% and 26.3% experienced mortality at U-M and BIDMC. The effect of sepsis on mortality was weakly correlated with sepsis risk (U-M: 0.35 [95% CI: 0.33-0.37], BIDMC: 0.31 [95% CI: 0.28-0.34]). High-risk patients identified by both stratification approaches overlapped by 66.8% and 52.8% at U-M and BIDMC, respectively. Accounting for risk of mortality identified an older population (U-M: age = 66.0 [interquartile range-IQR: 55.0-74.0] vs age = 63.0 [IQR: 51.0-72.0], BIDMC: age = 74.0 [IQR: 61.0-83.0] vs age = 68.0 [IQR: 59.0-78.0]). DISCUSSION: Predictive models that guide selective interventions ignore the effect of disease on downstream outcomes. Reformulating patient stratification to account for the estimated effect of disease on downstream outcomes identifies a different population compared to stratification on disease risk alone. CONCLUSION: Models that predict the risk of disease and ignore the effects of disease on downstream outcomes could be suboptimal for stratification.
Fahad Kamran, Donna Tjandra, Thomas S. Valley, Hallie C. Prescott, Nigam H. Shah, Vincent X. Liu, Eric Horvitz, Jenna Wiens
J. Am. Medical Informatics Assoc.5
2024 MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records
abstract
The ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care. However, evaluating LLMs on realistic text generation tasks for healthcare remains challenging. Existing question answering datasets for electronic health record (EHR) data fail to capture the complexity of information needs and documentation burdens experienced by clinicians. To address these challenges, we introduce MedAlign, a benchmark dataset of 983 natural language instructions for EHR data. MedAlign is curated by 15 clinicians (7 specialities), includes clinician-written reference responses for 303 instructions, and provides 276 longitudinal EHRs for grounding instruction-response pairs. We used MedAlign to evaluate 6 general domain LLMs, having clinicians rank the accuracy and quality of each LLM response. We found high error rates, ranging from 35% (GPT-4) to 68% (MPT-7B-Instruct), and 8.3% drop in accuracy moving from 32k to 2k context lengths for GPT-4. Finally, we report correlations between clinician rankings and automated natural language generation metrics as a way to rank LLMs without human review. MedAlign is provided under a research data use agreement to enable LLM evaluations on tasks aligned with clinician needs and preferences.
Scott L. Fleming, Alejandro Lozano, William J. Haberkorn, Jenelle A. Jindal, Eduardo Pontes Reis, Rahul Thapa, Louis Blankemeier, Julian Z. Genkins, Ethan Steinberg, Ashwin Nayak 0002, Birju Patel, Chia-Chun Chiang, Alison Callahan, Zepeng Huo, Sergios Gatidis, Scott J. Adams, Oluseyi Fayanju, Shreya J. Shah, Thomas Savage, Ethan Goh, Akshay Chaudhari, Nima Aghaeepour, Christopher D. Sharp, Michael A. Pfeffer, Percy Liang, Jonathan H. Chen, Keith E. Morse, Emma Brunskill, Jason Alan Fries, Nigam H. Shah
AAAI30
2024 MOTOR: A Time-to-Event Foundation Model For Structured Medical Records
abstract
We present a self-supervised, time-to-event (TTE) foundation model called MOTOR (Many Outcome Time Oriented Representations) which is pretrained on timestamped sequences of events in electronic health records (EHR) and health insurance claims. TTE models are used for estimating the probability distribution of the time until a specific event occurs, which is an important task in medical settings. TTE models provide many advantages over classification using fixed time horizons, including naturally handling censored observations, but are challenging to train with limited labeled data. MOTOR addresses this challenge by pretraining on up to 55M patient records (9B clinical events). We evaluate MOTOR's transfer learning performance on 19 tasks, across 3 patient databases (a private EHR system, MIMIC-IV, and Merative claims data). Task-specific models adapted from MOTOR improve time-dependent C statistics by 4.6\% over state-of-the-art, improve label efficiency by up to 95\%, and are more robust to temporal distributional shifts. We further evaluate cross-site portability by adapting our MOTOR foundation model for six prediction tasks on the MIMIC-IV dataset, where it outperforms all baselines. MOTOR is the first foundation model for medical TTE predictions and we release a 143M parameter pretrained model for research use at https://huggingface.co/StanfordShahLab/motor-t-base.
Ethan Steinberg, Jason Alan Fries, Yizhe Xu, Nigam H. Shah
ICLR4
2024 WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks
abstract
Existing ML benchmarks lack the depth and diversity of annotations needed for evaluating models on business process management (BPM) tasks. BPM is the practice of documenting, measuring, improving, and automating enterprise workflows. However, research has focused almost exclusively on one task -- full end-to-end automation using agents based on multimodal foundation models (FMs) like GPT-4. This focus on automation ignores the reality of how most BPM tools are applied today -- simply documenting the relevant workflow takes 60% of the time of the typical process optimization project. To address this gap we present WONDERBREAD, the first benchmark for evaluating multimodal FMs on BPM tasks beyond automation. Our contributions are: (1) a dataset containing 2928 documented workflow demonstrations; (2) 6 novel BPM tasks sourced from real-world applications ranging from workflow documentation to knowledge transfer to process improvement; and (3) an automated evaluation harness. Our benchmark shows that while state-of-the-art FMs can automatically generate documentation (e.g. recalling 88% of the steps taken in a video demonstration of a workflow), they struggle to re-apply that knowledge towards finer-grained validation of workflow completion (F1 < 0.3). We hope WONDERBREAD encourages the development of more "human-centered" AI tooling for enterprise applications and furthers the exploration of multimodal FMs for the broader universe of BPM tasks. We publish our dataset and experiments here: https://github.com/HazyResearch/wonderbread
Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan S. Khare, Tathagat Verma, Tibor Thompson, Miguel Angel Fuentes Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez 0003, Vardhan Agrawal, Althea Hudson, Nigam H. Shah, Christopher Ré
NeurIPS17
2024 Ensuring useful adoption of generative artificial intelligence in healthcare
abstract
OBJECTIVES: This article aims to examine how generative artificial intelligence (AI) can be adopted with the most value in health systems, in response to the Executive Order on AI. MATERIALS AND METHODS: We reviewed how technology has historically been deployed in healthcare, and evaluated recent examples of deployments of both traditional AI and generative AI (GenAI) with a lens on value. RESULTS: Traditional AI and GenAI are different technologies in terms of their capability and modes of current deployment, which have implications on value in health systems. DISCUSSION: Traditional AI when applied with a framework top-down can realize value in healthcare. GenAI in the short term when applied top-down has unclear value, but encouraging more bottom-up adoption has the potential to provide more benefit to health systems and patients. CONCLUSION: GenAI in healthcare can provide the most value for patients when health systems adapt culturally to grow with this new technology and its adoption patterns.
Jenelle A. Jindal, Matthew P. Lungren, Nigam H. Shah
J. Am. Medical Informatics Assoc.3
2024 Automating the Enterprise with Foundation Models
abstract
Automating enterprise workflows could unlock $4 trillion/year in productivity gains. Despite being of interest to the data management community for decades, the ultimate vision of end-to-end workflow automation has remained elusive. Current solutions rely on process mining and robotic process automation (RPA), in which a bot is hard-coded to follow a set of predefined rules for completing a workflow. Through case studies of a hospital and large B2B enterprise, we find that the adoption of RPA has been inhibited by high set-up costs (12--18 months), unreliable execution (60% initial accuracy), and burdensome maintenance (requiring multiple FTEs). Multimodal foundation models (FMs) such as GPT-4 offer a promising new approach for end-to-end workflow automation given their generalized reasoning and planning abilities. To study these capabilities we propose ECLAIR, a system to automate enterprise workflows with minimal human supervision. We conduct initial experiments showing that multimodal FMs can address the limitations of traditional RPA with (1) near-human-level understanding of workflows (93% accuracy on a workflow understanding task) and (2) instant set-up with minimal technical barrier (based solely on a natural language description of a workflow, ECLAIR achieves end-to-end completion rates of 40%). We identify human-AI collaboration, validation, and self-improvement as open challenges, and suggest ways they can be solved with data management techniques.
Michael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre, Nigam H. Shah, Christopher Ré
Proc. VLDB Endow.5
2023 INSPECT: A Multimodal Dataset for Patient Outcome Prediction of Pulmonary Embolisms
abstract
Synthesizing information from various data sources plays a crucial role in the practice of modern medicine. Current applications of artificial intelligence in medicine often focus on single-modality data due to a lack of publicly available, multimodal medical datasets. To address this limitation, we introduce INSPECT, which contains de-identified longitudinal records from a large cohort of pulmonary embolism (PE) patients, along with ground truth labels for multiple outcomes. INSPECT contains data from 19,402 patients, including CT images, sections of radiology reports, and structured electronic health record (EHR) data (including demographics, diagnoses, procedures, and vitals). Using our provided dataset, we develop and release a benchmark for evaluating several baseline modeling approaches on a variety of important PE related tasks. We evaluate image-only, EHR-only, and fused models. Trained models and the de-identified dataset are made available for non-commercial use under a data use agreement. To the best our knowledge, INSPECT is the largest multimodal dataset for enabling reproducible research on strategies for integrating 3D medical imaging and EHR data.
Shih-Cheng Huang, Zepeng Huo, Ethan Steinberg, Chia-Chun Chiang, Curt Langlotz, Matthew P. Lungren, Serena Yeung-Levy, Nigam H. Shah, Jason Alan Fries
NeurIPS8
2023 EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models
abstract
While the general machine learning (ML) community has benefited from public datasets, tasks, and models, the progress of ML in healthcare has been hampered by a lack of such shared assets. The success of foundation models creates new challenges for healthcare ML by requiring access to shared pretrained models to validate performance benefits. We help address these challenges through three contributions. First, we publish a new dataset, EHRSHOT, which contains de-identified structured data from the electronic health records (EHRs) of 6,739 patients from Stanford Medicine. Unlike MIMIC-III/IV and other popular EHR datasets, EHRSHOT is longitudinal and not restricted to ICU/ED patients. Second, we publish the weights of CLMBR-T-base, a 141M parameter clinical foundation model pretrained on the structured EHR data of 2.57M patients. We are one of the first to fully release such a model for coded EHR data; in contrast, most prior models released for clinical data (e.g. GatorTron, ClinicalBERT) only work with unstructured text and cannot process the rich, structured data within an EHR. We provide an end-to-end pipeline for the community to validate and build upon its performance. Third, we define 15 few-shot clinical prediction tasks, enabling evaluation of foundation models on benefits such as sample efficiency and task adaptation. Our model and dataset are available via a research data use agreement from here: https://stanfordaimi.azurewebsites.net/. Code to reproduce our results is available here: https://github.com/som-shahlab/ehrshot-benchmark.
Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason Alan Fries, Nigam H. Shah
NeurIPS5
2023 A framework to identify ethical concerns with ML-guided care workflows: a case study of mortality prediction to guide advance care planning
abstract
OBJECTIVE: Identifying ethical concerns with ML applications to healthcare (ML-HCA) before problems arise is now a stated goal of ML design oversight groups and regulatory agencies. Lack of accepted standard methodology for ethical analysis, however, presents challenges. In this case study, we evaluate use of a stakeholder "values-collision" approach to identify consequential ethical challenges associated with an ML-HCA for advanced care planning (ACP). Identification of ethical challenges could guide revision and improvement of the ML-HCA. MATERIALS AND METHODS: We conducted semistructured interviews of the designers, clinician-users, affiliated administrators, and patients, and inductive qualitative analysis of transcribed interviews using modified grounded theory. RESULTS: Seventeen stakeholders were interviewed. Five "values-collisions"-where stakeholders disagreed about decisions with ethical implications-were identified: (1) end-of-life workflow and how model output is introduced; (2) which stakeholders receive predictions; (3) benefit-harm trade-offs; (4) whether the ML design team has a fiduciary relationship to patients and clinicians; and, (5) how and if to protect early deployment research from external pressures, like news scrutiny, before research is completed. DISCUSSION: From these findings, the ML design team prioritized: (1) alternative workflow implementation strategies; (2) clarification that prediction was only evaluated for ACP need, not other mortality-related ends; and (3) shielding research from scrutiny until endpoint driven studies were completed. CONCLUSION: In this case study, our ethical analysis of this ML-HCA for ACP was able to identify multiple sites of intrastakeholder disagreement that mark areas of ethical and value tension. These findings provided a useful initial ethical screening.
Diana Cagliero, Natalie Deuitch, Nigam H. Shah, Chris Feudtner, Danton Char
J. Am. Medical Informatics Assoc.3
2023 DEPLOYR: a technical framework for deploying custom real-time machine learning models into the electronic medical record
abstract
OBJECTIVE: Heatlhcare institutions are establishing frameworks to govern and promote the implementation of accurate, actionable, and reliable machine learning models that integrate with clinical workflow. Such governance frameworks require an accompanying technical framework to deploy models in a resource efficient, safe and high-quality manner. Here we present DEPLOYR, a technical framework for enabling real-time deployment and monitoring of researcher-created models into a widely used electronic medical record system. MATERIALS AND METHODS: We discuss core functionality and design decisions, including mechanisms to trigger inference based on actions within electronic medical record software, modules that collect real-time data to make inferences, mechanisms that close-the-loop by displaying inferences back to end-users within their workflow, monitoring modules that track performance of deployed models over time, silent deployment capabilities, and mechanisms to prospectively evaluate a deployed model's impact. RESULTS: We demonstrate the use of DEPLOYR by silently deploying and prospectively evaluating 12 machine learning models trained using electronic medical record data that predict laboratory diagnostic results, triggered by clinician button-clicks in Stanford Health Care's electronic medical record. DISCUSSION: Our study highlights the need and feasibility for such silent deployment, because prospectively measured performance varies from retrospective estimates. When possible, we recommend using prospectively estimated performance measures during silent trials to make final go decisions for model deployment. CONCLUSION: Machine learning applications in healthcare are extensively researched, but successful translations to the bedside are rare. By describing DEPLOYR, we aim to inform machine learning deployment best practices and help bridge the model implementation gap.
Conor K. Corbin, Rob Maclay, Aakash Acharya, Sreedevi Mony, Soumya Punnathanam, Rahul Thapa, Nikesh Kotecha, Nigam H. Shah, Jonathan H. Chen
J. Am. Medical Informatics Assoc.8
2023 Self-supervised machine learning using adult inpatient data produces effective models for pediatric clinical prediction tasks
abstract
OBJECTIVE: Development of electronic health records (EHR)-based machine learning models for pediatric inpatients is challenged by limited training data. Self-supervised learning using adult data may be a promising approach to creating robust pediatric prediction models. The primary objective was to determine whether a self-supervised model trained in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients, for pediatric inpatient clinical prediction tasks. MATERIALS AND METHODS: This retrospective cohort study used EHR data and included patients with at least one admission to an inpatient unit. One admission per patient was randomly selected. Adult inpatients were 18 years or older while pediatric inpatients were more than 28 days and less than 18 years. Admissions were temporally split into training (January 1, 2008 to December 31, 2019), validation (January 1, 2020 to December 31, 2020), and test (January 1, 2021 to August 1, 2022) sets. Primary comparison was a self-supervised model trained in adult inpatients versus count-based logistic regression models trained in pediatric inpatients. Primary outcome was mean area-under-the-receiver-operating-characteristic-curve (AUROC) for 11 distinct clinical outcomes. Models were evaluated in pediatric inpatients. RESULTS: When evaluated in pediatric inpatients, mean AUROC of self-supervised model trained in adult inpatients (0.902) was noninferior to count-based logistic regression models trained in pediatric inpatients (0.868) (mean difference = 0.034, 95% CI=0.014-0.057; P < .001 for noninferiority and P = .006 for superiority). CONCLUSIONS: Self-supervised learning in adult inpatients was noninferior to logistic regression models trained in pediatric inpatients. This finding suggests transferability of self-supervised models trained in adult patients to pediatric patients, without requiring costly model retraining.
Joshua Lemmon, Lin Lawrence Guo, Ethan Steinberg, Keith E. Morse, Scott L. Fleming, Catherine Aftandilian, Stephen Pfohl, José D. Posada, Nigam H. Shah, Jason Alan Fries, Lillian Sung
J. Am. Medical Informatics Assoc.9
2023 Assessing the net benefit of machine learning models in the presence of resource constraints
abstract
OBJECTIVE: The objective of this study is to provide a method to calculate model performance measures in the presence of resource constraints, with a focus on net benefit (NB). MATERIALS AND METHODS: To quantify a model's clinical utility, the Equator Network's TRIPOD guidelines recommend the calculation of the NB, which reflects whether the benefits conferred by intervening on true positives outweigh the harms conferred by intervening on false positives. We refer to the NB achievable in the presence of resource constraints as the realized net benefit (RNB), and provide formulae for calculating the RNB. RESULTS: Using 4 case studies, we demonstrate the degree to which an absolute constraint (eg, only 3 available intensive care unit [ICU] beds) diminishes the RNB of a hypothetical ICU admission model. We show how the introduction of a relative constraint (eg, surgical beds that can be converted to ICU beds for very high-risk patients) allows us to recoup some of the RNB but with a higher penalty for false positives. DISCUSSION: RNB can be calculated in silico before the model's output is used to guide care. Accounting for the constraint changes the optimal strategy for ICU bed allocation. CONCLUSIONS: This study provides a method to account for resource constraints when planning model-based interventions, either to avoid implementations where constraints are expected to play a larger role or to design more creative solutions (eg, converted ICU beds) to overcome absolute constraints when possible.
Karandeep Singh, Nigam H. Shah, Andrew J. Vickers
J. Am. Medical Informatics Assoc.2
2023 Clinical utility gains from incorporating comorbidity and geographic location information into risk estimation equations for atherosclerotic cardiovascular disease
abstract
OBJECTIVE: There are over 363 customized risk models of the American College of Cardiology and the American Heart Association (ACC/AHA) pooled cohort equations (PCE) in the literature, but their gains in clinical utility are rarely evaluated. We build new risk models for patients with specific comorbidities and geographic locations and evaluate whether performance improvements translate to gains in clinical utility. MATERIALS AND METHODS: We retrain a baseline PCE using the ACC/AHA PCE variables and revise it to incorporate subject-level information of geographic location and 2 comorbidity conditions. We apply fixed effects, random effects, and extreme gradient boosting (XGB) models to handle the correlation and heterogeneity induced by locations. Models are trained using 2 464 522 claims records from Optum©'s Clinformatics® Data Mart and validated in the hold-out set (N = 1 056 224). We evaluate models' performance overall and across subgroups defined by the presence or absence of chronic kidney disease (CKD) or rheumatoid arthritis (RA) and geographic locations. We evaluate models' expected utility using net benefit and models' statistical properties using several discrimination and calibration metrics. RESULTS: The revised fixed effects and XGB models yielded improved discrimination, compared to baseline PCE, overall and in all comorbidity subgroups. XGB improved calibration for the subgroups with CKD or RA. However, the gains in net benefit are negligible, especially under low exchange rates. CONCLUSIONS: Common approaches to revising risk calculators incorporating extra information or applying flexible models may enhance statistical performance; however, such improvement does not necessarily translate to higher clinical utility. Thus, we recommend future works to quantify the consequences of using risk calculators to guide clinical decisions.
Yizhe Xu, Agata Foryciarz, Ethan Steinberg, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2023 APLUS: A Python library for usefulness simulations of machine learning models in healthcare
abstract
Despite the creation of thousands of machine learning (ML) models, the promise of improving patient care with ML remains largely unrealized. Adoption into clinical practice is lagging, in large part due to disconnects between how ML practitioners evaluate models and what is required for their successful integration into care delivery. Models are just one component of care delivery workflows whose constraints determine clinicians' abilities to act on models' outputs. However, methods to evaluate the usefulness of models in the context of their corresponding workflows are currently limited. To bridge this gap we developed APLUS, a reusable framework for quantitatively assessing via simulation the utility gained from integrating a model into a clinical workflow. We describe the APLUS simulation engine and workflow specification language, and apply it to evaluate a novel ML-based screening pathway for detecting peripheral artery disease at Stanford Health Care.
Michael Wornow, Elsie Gyang Ross, Alison Callahan, Nigam H. Shah
J. Biomed. Informatics4
2023 Principled estimation and evaluation of treatment effect heterogeneity: A case study application to dabigatran for patients with atrial fibrillation
abstract
OBJECTIVE: To apply the latest guidance for estimating and evaluating heterogeneous treatment effects (HTEs) in an end-to-end case study of the Long-term Anticoagulation Therapy (RE-LY) trial, and summarize the main takeaways from applying state-of-the-art metalearners and novel evaluation metrics in-depth to inform their applications to personalized care in biomedical research. METHODS: Based on the characteristics of the RE-LY data, we selected four metalearners (S-learner with Lasso, X-learner with Lasso, R-learner with random survival forest and Lasso, and causal survival forest) to estimate the HTEs of dabigatran. For the outcomes of (1) stroke or systemic embolism and (2) major bleeding, we compared dabigatran 150 mg, dabigatran 110 mg, and warfarin. We assessed the overestimation of treatment heterogeneity by the metalearners via a global null analysis and their discrimination and calibration ability using two novel metrics: rank-weighted average treatment effects (RATE) and estimated calibration error for treatment heterogeneity. Finally, we visualized the relationships between estimated treatment effects and baseline covariates using partial dependence plots. RESULTS: The RATE metric suggested that either the applied metalearners had poor performance of estimating HTEs or there was no treatment heterogeneity for either the stroke/SE or major bleeding outcome of any treatment comparison. Partial dependence plots revealed that several covariates had consistent relationships with the treatment effects estimated by multiple metalearners. The applied metalearners showed differential performance across outcomes and treatment comparisons, and the X- and R-learners yielded smaller calibration errors than the others. CONCLUSIONS: HTE estimation is difficult, and a principled estimation and evaluation process is necessary to provide reliable evidence and prevent false discoveries. We have demonstrated how to choose appropriate metalearners based on specific data properties, applied them using the off-the-shelf implementation tool survlearners, and evaluated their performance using recently defined formal metrics. We suggest that clinical implications should be drawn based on the common trends across the applied metalearners.
Yizhe Xu, Katelyn K. Bechler, Alison Callahan, Nigam H. Shah
J. Biomed. Informatics4
2022 Predicting patients who are likely to develop Lupus Nephritis of those newly diagnosed with Systemic Lupus Erythematosus
Katelyn K. Bechler, Nigam H. Shah
AMIA2
2021 Multi-Modal Data Science for Healthcare: State of the Art, Challenges, and Opportunities
Yuan Luo 0001, Fei Wang 0001, Benjamin S. Glicksberg, Jessilyn Dunn, Nigam H. Shah
AMIA5
2021 Computational drug repositioning of atorvastatin for ulcerative colitis
abstract
OBJECTIVE: Ulcerative colitis (UC) is a chronic inflammatory disorder with limited effective therapeutic options for long-term treatment and disease maintenance. We hypothesized that a multi-cohort analysis of independent cohorts representing real-world heterogeneity of UC would identify a robust transcriptomic signature to improve identification of FDA-approved drugs that can be repurposed to treat patients with UC. MATERIALS AND METHODS: We performed a multi-cohort analysis of 272 colon biopsy transcriptome samples across 11 publicly available datasets to identify a robust UC disease gene signature. We compared the gene signature to in vitro transcriptomic profiles induced by 781 FDA-approved drugs to identify potential drug targets. We used a retrospective cohort study design modeled after a target trial to evaluate the protective effect of predicted drugs on colectomy risk in patients with UC from the Stanford Research Repository (STARR) database and Optum Clinformatics DataMart. RESULTS: Atorvastatin treatment had the highest inverse-correlation with the UC gene signature among non-oncolytic FDA-approved therapies. In both STARR (n = 827) and Optum (n = 7821), atorvastatin intake was significantly associated with a decreased risk of colectomy, a marker of treatment-refractory disease, compared to patients prescribed a comparator drug (STARR: HR = 0.47, P = .03; Optum: HR = 0.66, P = .03), irrespective of age and length of atorvastatin treatment. DISCUSSION & CONCLUSION: These findings suggest that atorvastatin may serve as a novel therapeutic option for ameliorating disease in patients with UC. Importantly, we provide a systematic framework for integrating publicly available heterogeneous molecular data with clinical data at a large scale to repurpose existing FDA-approved drugs for a wide range of human diseases.
Lawrence Bai, Madeleine K. D. Scott, Ethan Steinberg, Laurynas Kalesinskas, Aida Habtezion, Nigam H. Shah, Purvesh Khatri
J. Am. Medical Informatics Assoc.6
2021 ACE: the Advanced Cohort Engine for searching longitudinal patient records
abstract
OBJECTIVE: To propose a paradigm for a scalable time-aware clinical data search, and to describe the design, implementation and use of a search engine realizing this paradigm. MATERIALS AND METHODS: The Advanced Cohort Engine (ACE) uses a temporal query language and in-memory datastore of patient objects to provide a fast, scalable, and expressive time-aware search. ACE accepts data in the Observational Medicine Outcomes Partnership Common Data Model, and is configurable to balance performance with compute cost. ACE's temporal query language supports automatic query expansion using clinical knowledge graphs. The ACE API can be used with R, Python, Java, HTTP, and a Web UI. RESULTS: ACE offers an expressive query language for complex temporal search across many clinical data types with multiple output options. ACE enables electronic phenotyping and cohort-building with subsecond response times in searching the data of millions of patients for a variety of use cases. DISCUSSION: ACE enables fast, time-aware search using a patient object-centric datastore, thereby overcoming many technical and design shortcomings of relational algebra-based querying. Integrating electronic phenotype development with cohort-building enables a variety of high-value uses for a learning health system. Tradeoffs include the need to learn a new query language and the technical setup burden. CONCLUSION: ACE is a tool that combines a unique query language for time-aware search of longitudinal patient records with a patient object datastore for rapid electronic phenotyping, cohort extraction, and exploratory data analyses.
Alison Callahan, Vladimir Polony, José D. Posada, Juan M. Banda, Saurabh Gombar, Nigam H. Shah
J. Am. Medical Informatics Assoc.6
2021 Automated model versus treating physician for predicting survival time of patients with metastatic cancer
abstract
OBJECTIVE: Being able to predict a patient's life expectancy can help doctors and patients prioritize treatments and supportive care. For predicting life expectancy, physicians have been shown to outperform traditional models that use only a few predictor variables. It is possible that a machine learning model that uses many predictor variables and diverse data sources from the electronic medical record can improve on physicians' performance. For patients with metastatic cancer, we compared accuracy of life expectancy predictions by the treating physician, a machine learning model, and a traditional model. MATERIALS AND METHODS: A machine learning model was trained using 14 600 metastatic cancer patients' data to predict each patient's distribution of survival time. Data sources included note text, laboratory values, and vital signs. From 2015-2016, 899 patients receiving radiotherapy for metastatic cancer were enrolled in a study in which their radiation oncologist estimated life expectancy. Survival predictions were also made by the machine learning model and a traditional model using only performance status. Performance was assessed with area under the curve for 1-year survival and calibration plots. RESULTS: The radiotherapy study included 1190 treatment courses in 899 patients. A total of 879 treatment courses in 685 patients were included in this analysis. Median overall survival was 11.7 months. Physicians, machine learning model, and traditional model had area under the curve for 1-year survival of 0.72 (95% CI 0.63-0.81), 0.77 (0.73-0.81), and 0.68 (0.65-0.71), respectively. CONCLUSIONS: The machine learning model's predictions were more accurate than those of the treating physician or a traditional model.
Michael Gensheimer 0001, Sonya Aggarwal, Kathryn R. K. Benson, Justin N. Carter, Solomon Henry, Douglas J. Wood, Scott G. Soltys, Steven Hancock, Erqi Pollom, Nigam H. Shah, Daniel T. Chang
J. Am. Medical Informatics Assoc.10
2021 Conflicting information from the Food and Drug Administration: Missed opportunity to lead standards for safe and effective medical artificial intelligence solutions
abstract
The Food & Drug Administration (FDA) is considering the permanent exemption of premarket notification requirements for several Class I and II medical device products, including several artificial Intelligence (AI)-driven devices. The exemption is based on the need to rapidly more quickly disseminate devices to the public, estimated cost-savings, a lack of documented adverse events reported to the FDA's database. However, this ignores emerging issues related to AI-based devices, including utility, reproducibility and bias that may not only affect an individual but entire populations. We urge the FDA to reinforce the messaging on safety and effectiveness regulations of AI-based Software as a Medical Device products to better promote fair AI-driven clinical decision tools and for preventing harm to the patients we serve.
Tina Hernandez-Boussard, Matthew P. Lungren, Nigam H. Shah
J. Am. Medical Informatics Assoc.3
2021 Corrigendum: Conflicting information from the Food and Drug Administration: Missed opportunity to lead standards for safe and effective medical artificial intelligence solutions
abstract
Journal of the American Medical Informatics Association, 2021, doi: 10.1093/jamia/ocab035 The author name “Matthew P Lungren” was incorrectly given as “Matthew P Lundgren”. “CDRH” should have been defined at its first appearance, and incorrectly appeared in the second paragraph as “CDHR”. These errors have been corrected online.
Tina Hernandez-Boussard, Matthew P. Lungren, Nigam H. Shah
J. Am. Medical Informatics Assoc.3
2021 A framework for making predictive models useful in practice
abstract
OBJECTIVE: To analyze the impact of factors in healthcare delivery on the net benefit of triggering an Advanced Care Planning (ACP) workflow based on predictions of 12-month mortality. MATERIALS AND METHODS: We built a predictive model of 12-month mortality using electronic health record data and evaluated the impact of healthcare delivery factors on the net benefit of triggering an ACP workflow based on the models' predictions. Factors included nonclinical reasons that make ACP inappropriate: limited capacity for ACP, inability to follow up due to patient discharge, and availability of an outpatient workflow to follow up on missed cases. We also quantified the relative benefits of increasing capacity for inpatient ACP versus outpatient ACP. RESULTS: Work capacity constraints and discharge timing can significantly reduce the net benefit of triggering the ACP workflow based on a model's predictions. However, the reduction can be mitigated by creating an outpatient ACP workflow. Given limited resources to either add capacity for inpatient ACP versus developing outpatient ACP capability, the latter is likely to provide more benefit to patient care. DISCUSSION: The benefit of using a predictive model for identifying patients for interventions is highly dependent on the capacity to execute the workflow triggered by the model. We provide a framework for quantifying the impact of healthcare delivery factors and work capacity constraints on achieved benefit. CONCLUSION: An analysis of the sensitivity of the net benefit realized by a predictive model triggered clinical workflow to various healthcare delivery factors is necessary for making predictive models useful in practice.
Kenneth Jung, Sehj Kashyap, Anand Avati, Stephanie Harman, Heather Shaw, Ron C. Li, Margaret Smith, Kenny Shum, Jacob Javitz, Yohan Vetteth, Tina Seto, Steven C. Bagley, Nigam H. Shah
J. Am. Medical Informatics Assoc.13
2021 A survey of extant organizational and computational setups for deploying predictive models in health systems
abstract
OBJECTIVE: Artificial intelligence (AI) and machine learning (ML) enabled healthcare is now feasible for many health systems, yet little is known about effective strategies of system architecture and governance mechanisms for implementation. Our objective was to identify the different computational and organizational setups that early-adopter health systems have utilized to integrate AI/ML clinical decision support (AI-CDS) and scrutinize their trade-offs. MATERIALS AND METHODS: We conducted structured interviews with health systems with AI deployment experience about their organizational and computational setups for deploying AI-CDS at point of care. RESULTS: We contacted 34 health systems and interviewed 20 healthcare sites (58% response rate). Twelve (60%) sites used the native electronic health record vendor configuration for model development and deployment, making it the most common shared infrastructure. Nine (45%) sites used alternative computational configurations which varied significantly. Organizational configurations for managing AI-CDS were distinguished by how they identified model needs, built and implemented models, and were separable into 3 major types: Decentralized translation (n = 10, 50%), IT Department led (n = 2, 10%), and AI in Healthcare (AIHC) Team (n = 8, 40%). DISCUSSION: No singular computational configuration enables all current use cases for AI-CDS. Health systems need to consider their desired applications for AI-CDS and whether investment in extending the off-the-shelf infrastructure is needed. Each organizational setup confers trade-offs for health systems planning strategies to implement AI-CDS. CONCLUSION: Health systems will be able to use this framework to understand strengths and weaknesses of alternative organizational and computational setups when designing their strategy for artificial intelligence.
Sehj Kashyap, Keith E. Morse, Birju Patel, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2021 Learning decision thresholds for risk stratification models from aggregate clinician behavior
abstract
Using a risk stratification model to guide clinical practice often requires the choice of a cutoff-called the decision threshold-on the model's output to trigger a subsequent action such as an electronic alert. Choosing this cutoff is not always straightforward. We propose a flexible approach that leverages the collective information in treatment decisions made in real life to learn reference decision thresholds from physician practice. Using the example of prescribing a statin for primary prevention of cardiovascular disease based on 10-year risk calculated by the 2013 pooled cohort equations, we demonstrate the feasibility of using real-world data to learn the implicit decision threshold that reflects existing physician behavior. Learning a decision threshold in this manner allows for evaluation of a proposed operating point against the threshold reflective of the community standard of care. Furthermore, this approach can be used to monitor and audit model-guided clinical decision making following model deployment.
Birju Patel, Ethan Steinberg, Stephen Pfohl, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2021 Improving hospital readmission prediction using individualized utility analysis
abstract
OBJECTIVE: Machine learning (ML) models for allocating readmission-mitigating interventions are typically selected according to their discriminative ability, which may not necessarily translate into utility in allocation of resources. Our objective was to determine whether ML models for allocating readmission-mitigating interventions have different usefulness based on their overall utility and discriminative ability. MATERIALS AND METHODS: We conducted a retrospective utility analysis of ML models using claims data acquired from the Optum Clinformatics Data Mart, including 513,495 commercially-insured inpatients (mean [SD] age 69 [19] years; 294,895 [57%] Female) over the period January 2016 through January 2017 from all 50 states with mean 90 day cost of $11,552. Utility analysis estimates the cost, in dollars, of allocating interventions for lowering readmission risk based on the reduction in the 90-day cost. RESULTS: Allocating readmission-mitigating interventions based on a GBDT model trained to predict readmissions achieved an estimated utility gain of $104 per patient, and an AUC of 0.76 (95% CI 0.76, 0.77); allocating interventions based on a model trained to predict cost as a proxy achieved a higher utility of $175.94 per patient, and an AUC of 0.62 (95% CI 0.61, 0.62). A hybrid model combining both intervention strategies is comparable with the best models on either metric. Estimated utility varies by intervention cost and efficacy, with each model performing the best under different intervention settings. CONCLUSION: We demonstrate that machine learning models may be ranked differently based on overall utility and discriminative ability. Machine learning models for allocation of limited health resources should consider directly optimizing for utility.
Michael Ko, Emma Chen, Ashwin Agrawal, Pranav Rajpurkar, Anand Avati, Andrew Y. Ng, Sanjay Basu, Nigam H. Shah
J. Biomed. Informatics8
2021 An empirical characterization of fair machine learning for clinical risk prediction
Stephen Pfohl, Agata Foryciarz, Nigam H. Shah
J. Biomed. Informatics3
2021 Language models are an effective representation learning technique for electronic health record data
Ethan Steinberg, Kenneth Jung, Jason Alan Fries, Conor K. Corbin, Stephen Pfohl, Nigam H. Shah
J. Biomed. Informatics6
2021 Summarizing Patients Like Mine via an On-demand Consultation Service
Nigam H. Shah
Proc. VLDB Endow.1
2020 Data Quality Assessment of Laboratory Data
Vojtech Huser, Clair Blacketer, Karthik Natarajan, Robert T. Miller, Andrew E. Williams, Selva Muthu Kumaran Sathappan, José D. Posada, Nigam H. Shah
AMIA8
2020 Normalizing Clinical Document Titles to LOINC Document Ontology: an Initial Study
Xu Zuo, Jianfu Li, Bo Zhao 0001, Yujia Zhou 0003, Jon D. Duke, Karthik Natarajan, George Hripcsak, Nigam H. Shah, Juan M. Banda, Ruth M. Reeves, Hua Xu 0001
AMIA9
2020 MINIMAR (MINimum Information for Medical AI Reporting): Developing reporting standards for artificial intelligence in health care
abstract
The rise of digital data and computing power have contributed to significant advancements in artificial intelligence (AI), leading to the use of classification and prediction models in health care to enhance clinical decision-making for diagnosis, treatment and prognosis. However, such advances are limited by the lack of reporting standards for the data used to develop those models, the model architecture, and the model evaluation and validation processes. Here, we present MINIMAR (MINimum Information for Medical AI Reporting), a proposal describing the minimum information necessary to understand intended predictions, target populations, and hidden biases, and the ability to generalize these emerging technologies. We call for a standard to accurately and responsibly report on AI in health care. This will facilitate the design and implementation of these models and promote the development and use of associated clinical decision support tools, as well as manage concerns regarding accuracy and bias.
Tina Hernandez-Boussard, Selen Bozkurt, John P. A. Ioannidis, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2020 Measure what matters: Counts of hospitalized patients are a better metric for health system capacity planning for a reopening
abstract
OBJECTIVE: Responding to the COVID-19 pandemic requires accurate forecasting of health system capacity requirements using readily available inputs. We examined whether testing and hospitalization data could help quantify the anticipated burden on the health system given shelter-in-place (SIP) order. MATERIALS AND METHODS: 16,103 SARS-CoV-2 RT-PCR tests were performed on 15,807 patients at Stanford facilities between March 2 and April 11, 2020. We analyzed the fraction of tested patients that were confirmed positive for COVID-19, the fraction of those needing hospitalization, and the fraction requiring ICU admission over the 40 days between March 2nd and April 11th 2020. RESULTS: We find a marked slowdown in the hospitalization rate within ten days of SIP even as cases continued to rise. We also find a shift towards younger patients in the age distribution of those testing positive for COVID-19 over the four weeks of SIP. The impact of this shift is a divergence between increasing positive case confirmations and slowing new hospitalizations, both of which affects the demand on health systems. CONCLUSION: Without using local hospitalization rates and the age distribution of positive patients, current models are likely to overestimate the resource burden of COVID-19. It is imperative that health systems start using these data to quantify effects of SIP and aid reopening planning.
Sehj Kashyap, Saurabh Gombar, Steve Yadlowsky, Alison Callahan, Jason Alan Fries, Benjamin A. Pinsky, Nigam H. Shah
J. Am. Medical Informatics Assoc.7
2020 Development and validation of phenotype classifiers across multiple sites in the observational health data sciences and informatics network
abstract
OBJECTIVE: Accurate electronic phenotyping is essential to support collaborative observational research. Supervised machine learning methods can be used to train phenotype classifiers in a high-throughput manner using imperfectly labeled data. We developed 10 phenotype classifiers using this approach and evaluated performance across multiple sites within the Observational Health Data Sciences and Informatics (OHDSI) network. MATERIALS AND METHODS: We constructed classifiers using the Automated PHenotype Routine for Observational Definition, Identification, Training and Evaluation (APHRODITE) R-package, an open-source framework for learning phenotype classifiers using datasets in the Observational Medical Outcomes Partnership Common Data Model. We labeled training data based on the presence of multiple mentions of disease-specific codes. Performance was evaluated on cohorts derived using rule-based definitions and real-world disease prevalence. Classifiers were developed and evaluated across 3 medical centers, including 1 international site. RESULTS: Compared to the multiple mentions labeling heuristic, classifiers showed a mean recall boost of 0.43 with a mean precision loss of 0.17. Performance decreased slightly when classifiers were shared across medical centers, with mean recall and precision decreasing by 0.08 and 0.01, respectively, at a site within the USA, and by 0.18 and 0.10, respectively, at an international site. DISCUSSION AND CONCLUSION: We demonstrate a high-throughput pipeline for constructing and sharing phenotype classifiers across sites within the OHDSI network using APHRODITE. Classifiers exhibit good portability between sites within the USA, however limited portability internationally, indicating that classifier generalizability may have geographic limitations, and, consequently, sharing the classifier-building recipe, rather than the pretrained classifiers, may be more useful for facilitating collaborative observational research.
Mehr Kashyap, Martin G. Seneviratne, Juan M. Banda, Thomas Falconer, Borim Ryu, Sooyoung Yoo, George Hripcsak, Nigam H. Shah
J. Am. Medical Informatics Assoc.8
2020 Deep phenotyping: Embracing complexity and temporality - Towards scalability, portability, and interoperability
Chunhua Weng, Nigam H. Shah, George Hripcsak
J. Biomed. Informatics2
2019 Creating Fair Models of Atherosclerotic Cardiovascular Disease Risk
abstract
Guidelines for the management of atherosclerotic cardiovascular disease (ASCVD) recommend the use of risk stratification models to identify patients most likely to benefit from cholesterol-lowering and other therapies. These models have differential performance across race and gender groups with inconsistent behavior across studies, potentially resulting in an inequitable distribution of beneficial therapy. In this work, we leverage adversarial learning and a large observational cohort extracted from electronic health records (EHRs) to develop a "fair" ASCVD risk prediction model with reduced variability in error rates across groups. We empirically demonstrate that our approach is capable of aligning the distribution of risk predictions conditioned on the outcome across several groups simultaneously for models built from high-dimensional EHR data. We also discuss the relevance of these results in the context of the empirical trade-off between fairness and model performance.
Stephen Pfohl, Ben J. Marafino, Adrien Coulet, Fátima Rodriguez, Latha Palaniappan, Nigam H. Shah
AIES6
2019 Countdown Regression: Sharp and Calibrated Survival Predictions
Anand Avati, Tony Duan, Sharon Zhou, Kenneth Jung, Nigam H. Shah, Andrew Y. Ng
UAI5
2019 The number needed to benefit: estimating the value of predictive analytics in healthcare
abstract
Predictive analytics in health care has generated increasing enthusiasm recently, as reflected in a rapidly growing body of predictive models reported in literature and in real-time embedded models using electronic health record data. However, estimating the benefit of applying any single model to a specific clinical problem remains challenging today. Developing a shared framework for estimating model value is therefore critical to facilitate the effective, safe, and sustainable use of predictive tools into the future. We highlight key concepts within the prediction-action dyad that together are expected to impact model benefit. These include factors relevant to model prediction (including the number needed to screen) as well as those relevant to the subsequent action (number needed to treat). In the simplest terms, a number needed to benefit contextualizes the numbers needed to screen and treat, offering an opportunity to estimate the value of a clinical predictive model in action.
Vincent X. Liu, David W. Bates, Jenna Wiens, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2019 Predicting need for advanced illness or palliative care in a primary care population using electronic health record data
Kenneth Jung, Sylvia E. K. Sudat, Nicole Kwon, Walter F. Stewart, Nigam H. Shah
J. Biomed. Informatics5
2018 Treatment Pathways in Patients with Cancer Using a Large-scale Observational Data Network
Patrick B. Ryan, Karthik Natarajan, Thomas Falconer, Christian G. Reich, Rohit Vashisht, Nigam H. Shah, George Hripcsak
AMIA7
2018 Transfer learning to adapt predictive models for pediatric patients in the EHR
Stephen Pfohl, Nigam H. Shah
AMIA2
2018 Identifying cases of metastatic prostate cancer using machine learning on electronic health records
Martin G. Seneviratne, Juan M. Banda, James D. Brooks, Nigam H. Shah, Tina Hernandez-Boussard
AMIA4
2018 An evaluation of clinical order patterns machine-learned from clinician cohorts stratified by patient mortality outcomes
abstract
Evaluate the quality of clinical order practice patterns machine-learned from clinician cohorts stratified by patient mortality outcomes. Inpatient electronic health records from 2010 to 2013 were extracted from a tertiary academic hospital. Clinicians (n = 1822) were stratified into low-mortality (21.8%, n = 397) and high-mortality (6.0%, n = 110) extremes using a two-sided P-value score quantifying deviation of observed vs. expected 30-day patient mortality rates. Three patient cohorts were assembled: patients seen by low-mortality clinicians, high-mortality clinicians, and an unfiltered crowd of all clinicians (n = 1046, 1046, and 5230 post-propensity score matching, respectively). Predicted order lists were automatically generated from recommender system algorithms trained on each patient cohort and evaluated against (i) real-world practice patterns reflected in patient cases with better-than-expected mortality outcomes and (ii) reference standards derived from clinical practice guidelines. Across six common admission diagnoses, order lists learned from the crowd demonstrated the greatest alignment with guideline references (AUROC range = 0.86–0.91), performing on par or better than those learned from low-mortality clinicians (0.79–0.84, P < 10−5) or manually-authored hospital order sets (0.65–0.77, P < 10−3). The same trend was observed in evaluating model predictions against better-than-expected patient cases, with the crowd model (AUROC mean = 0.91) outperforming the low-mortality model (0.87, P < 10−16) and order set benchmarks (0.78, P < 10−35). Whether machine-learning models are trained on all clinicians or a subset of experts illustrates a bias-variance tradeoff in data usage. Defining robust metrics to assess quality based on internal (e.g. practice patterns from better-than-expected patient cases) or external reference standards (e.g. clinical practice guidelines) is critical to assess decision support content. Learning relevant decision support content from all clinicians is as, if not more, robust than learning from a select subgroup of clinicians favored by patient outcomes.
Jason K. Wang, Jason Hom, Santhosh Balasubramanian, Alejandro Schuler, Nigam H. Shah, Mary K. Goldstein, Michael T. M. Baiocchi, Jonathan H. Chen
J. Biomed. Informatics5
2018 Call for papers: Deep phenotyping for Precision Medicine
Chunhua Weng, Nigam H. Shah, George Hripcsak
J. Biomed. Informatics2
2017 From Large-Scale Network Analytics to Clinical Solutions in OHDSI
Jon D. Duke, George Hripcsak, Patrick B. Ryan, Nigam H. Shah
AMIA4
2017 Impact of Clinician Experience on Machine Learned Clinical Order Patterns
Jason K. Wang, Alejandro Schuler, Nigam H. Shah, Jonathan H. Chen
AMIA3
2017 Improving palliative care with deep learning
abstract
BACKGROUND: Access to palliative care is a key quality metric which most healthcare organizations strive to improve. The primary challenges to increasing palliative care access are a combination of physicians over-estimating patient prognoses, and a shortage of palliative staff in general. This, in combination with treatment inertia can result in a mismatch between patient wishes, and their actual care towards the end of life. METHODS: In this work, we address this problem, with Institutional Review Board approval, using machine learning and Electronic Health Record (EHR) data of patients. We train a Deep Neural Network model on the EHR data of patients from previous years, to predict mortality of patients within the next 3-12 month period. This prediction is used as a proxy decision for identifying patients who could benefit from palliative care. RESULTS: The EHR data of all admitted patients are evaluated every night by this algorithm, and the palliative care team is automatically notified of the list of patients with a positive prediction. In addition, we present a novel technique for decision interpretation, using which we provide explanations for the model's predictions. CONCLUSION: The automatic screening and notification saves the palliative care team the burden of time consuming chart reviews of all patients, and allows them to take a proactive approach in reaching out to such patients rather then relying on referrals from the treating physicians.
Anand Avati, Kenneth Jung, Stephanie Harman, Lance Downing, Andrew Y. Ng, Nigam H. Shah
BIBM6
2017 Synergistic drug combinations from electronic health records and gene expression
abstract
OBJECTIVE: Using electronic health records (EHRs) and biomolecular data, we sought to discover drug pairs with synergistic repurposing potential. EHRs provide real-world treatment and outcome patterns, while complementary biomolecular data, including disease-specific gene expression and drug-protein interactions, provide mechanistic understanding. METHOD: We applied Group Lasso INTERaction NETwork (glinternet), an overlap group lasso penalty on a logistic regression model, with pairwise interactions to identify variables and interacting drug pairs associated with reduced 5-year mortality using EHRs of 9945 breast cancer patients. We identified differentially expressed genes from 14 case-control human breast cancer gene expression datasets and integrated them with drug-protein networks. Drugs in the network were scored according to their association with breast cancer individually or in pairs. Lastly, we determined whether synergistic drug pairs found in the EHRs were enriched among synergistic drug pairs from gene-expression data using a method similar to gene set enrichment analysis. RESULTS: From EHRs, we discovered 3 drug-class pairs associated with lower mortality: anti-inflammatories and hormone antagonists, anti-inflammatories and lipid modifiers, and lipid modifiers and obstructive airway drugs. The first 2 pairs were also enriched among pairs discovered using gene expression data and are supported by molecular interactions in drug-protein networks and preclinical and epidemiologic evidence. CONCLUSIONS: This is a proof-of-concept study demonstrating that a combination of complementary data sources, such as EHRs and gene expression, can corroborate discoveries and provide mechanistic insight into drug synergism for repurposing.
Yen S. Low, Aaron C. Daugherty, Elizabeth A. Schroeder, Tina Seto, Susan C. Weber, Michael Lim, Trevor J. Hastie, Maya Mathur, Manisha Desai, Carl Farrington, Andrew A. Radin, Marina Sirota, Pragati Kenkare, Caroline A. Thompson, Peter P. Yu, Scarlett L. Gomez, George W. Sledge, Allison W. Kurian, Nigam H. Shah
J. Am. Medical Informatics Assoc.20
2017 Toward multimodal signal detection of adverse drug reactions
Rave Harpaz, William DuMouchel, Martijn J. Schuemie, Olivier Bodenreider, Carol Friedman, Eric Horvitz, Anna Ripple, Alfred Sorbello, Ryen W. White, Rainer Winnenburg, Nigam H. Shah
J. Biomed. Informatics11
2016 Ensuring Reproducibility in Observational Research: Building and Sharing Knowledge Resources in the OHDSI Network
Jon D. Duke, Nigam H. Shah, George Hripcsak, Patrick B. Ryan
AMIA2
2016 Big Data for Healthcare and Life Sciences: Learning Useful Insights from Imperfect Data
Jianying Hu, Nigam H. Shah, Bradley A. Malin, Patrick B. Ryan
AMIA2
2016 Learning Effective Treatment Pathways for Type-2 Diabetes from a clinical data warehouse
Rohit Vashisht, Kenneth Jung, Nigam H. Shah
AMIA3
2016 The digital revolution in phenotyping
abstract
Phenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data.
Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier
Briefings Bioinform.5
2016 Generalized enrichment analysis improves the detection of adverse drug events from the biomedical literature
abstract
BACKGROUND: Identification of associations between marketed drugs and adverse events from the biomedical literature assists drug safety monitoring efforts. Assessing the significance of such literature-derived associations and determining the granularity at which they should be captured remains a challenge. Here, we assess how defining a selection of adverse event terms from MeSH, based on information content, can improve the detection of adverse events for drugs and drug classes. RESULTS: We analyze a set of 105,354 candidate drug adverse event pairs extracted from article indexes in MEDLINE. First, we harmonize extracted adverse event terms by aggregating them into higher-level MeSH terms based on the terms' information content. Then, we determine statistical enrichment of adverse events associated with drug and drug classes using a conditional hypergeometric test that adjusts for dependencies among associated terms. We compare our results with methods based on disproportionality analysis (proportional reporting ratio, PRR) and quantify the improvement in signal detection with our generalized enrichment analysis (GEA) approach using a gold standard of drug-adverse event associations spanning 174 drugs and four events. For single drugs, the best GEA method (Precision: .92/Recall: .71/F1-measure: .80) outperforms the best PRR based method (.69/.69/.69) on all four adverse event outcomes in our gold standard. For drug classes, our GEA performs similarly (.85/.69/.74) when increasing the level of abstraction for adverse event terms. Finally, on examining the 1609 individual drugs in our MEDLINE set, which map to chemical substances in ATC, we find signals for 1379 drugs (10,122 unique adverse event associations) on applying GEA with p < 0.005. CONCLUSIONS: We present an approach based on generalized enrichment analysis that can be used to detect associations between drugs, drug classes and adverse events at a given level of granularity, at the same time correcting for known dependencies among events. Our study demonstrates the use of GEA, and the importance of choosing appropriate abstraction levels to complement current drug safety methods. We provide an R package for exploration of alternative abstraction levels of adverse event terms based on information content.
Rainer Winnenburg, Nigam H. Shah
BMC Bioinform.2
2016 Learning statistical models of phenotypes using noisy labeled training data
abstract
OBJECTIVE: Traditionally, patient groups with a phenotype are selected through rule-based definitions whose creation and validation are time-consuming. Machine learning approaches to electronic phenotyping are limited by the paucity of labeled training datasets. We demonstrate the feasibility of utilizing semi-automatically labeled training sets to create phenotype models via machine learning, using a comprehensive representation of the patient medical record. METHODS: We use a list of keywords specific to the phenotype of interest to generate noisy labeled training data. We train L1 penalized logistic regression models for a chronic and an acute disease and evaluate the performance of the models against a gold standard. RESULTS: Our models for Type 2 diabetes mellitus and myocardial infarction achieve precision and accuracy of 0.90, 0.89, and 0.86, 0.89, respectively. Local implementations of the previously validated rule-based definitions for Type 2 diabetes mellitus and myocardial infarction achieve precision and accuracy of 0.96, 0.92 and 0.84, 0.87, respectively.We have demonstrated feasibility of learning phenotype models using imperfectly labeled data for a chronic and acute phenotype. Further research in feature engineering and in specification of the keyword list can improve the performance of the models and the scalability of the approach. CONCLUSIONS: Our method provides an alternative to manual labeling for creating training sets for statistical models of phenotypes. Such an approach can accelerate research with large observational healthcare datasets and may also be used to create local phenotype models.
Vibhu Agarwal, Tanya Podchiyska, Juan M. Banda, Veena Goel, Tiffany I. Leung, Evan P. Minty, Timothy E. Sweeney, Elsie Gyang, Nigam H. Shah
J. Am. Medical Informatics Assoc.9
2016 Harnessing next-generation informatics for personalizing medicine: a report from AMIA's 2014 Health Policy Invitational Meeting
abstract
The American Medical Informatics Association convened the 2014 Health Policy Invitational Meeting to develop recommendations for updates to current policies and to establish an informatics research agenda for personalizing medicine. In particular, the meeting focused on discussing informatics challenges related to personalizing care through the integration of genomic or other high-volume biomolecular data with data from clinical systems to make health care more efficient and effective. This report summarizes the findings (n = 6) and recommendations (n = 15) from the policy meeting, which were clustered into 3 broad areas: (1) policies governing data access for research and personalization of care; (2) policy and research needs for evolving data interpretation and knowledge representation; and (3) policy and research needs to ensure data integrity and preservation. The meeting outcome underscored the need to address a number of important policy and technical considerations in order to realize the potential of personalized or precision medicine in actual clinical contexts.
Laura K. Wiley, Peter Tarczy-Hornoch, Joshua C. Denny, Robert R. Freimuth, Casey Overby Taylor, Nigam H. Shah, Ross D. Martin, Indra Neil Sarkar
J. Am. Medical Informatics Assoc.6
2016 An unsupervised learning method to identify reference intervals from a clinical database
Sarah Poole, Lee Frederick Schroeder, Nigam H. Shah
J. Biomed. Informatics3
2015 Recent Advances in Computational Drug Repositioning
Atul J. Butte, Nigam H. Shah, Nicholas P. Tatonetti, Hua Xu 0001
AMIA2
2015 The Value of an Open-Source Observational Research Collaboratory: Results from the OHDSI Initiative
Jon D. Duke, George Hripcsak, Nigam H. Shah, Patrick B. Ryan
AMIA3
2015 Provenance-Centered Dataset of Drug-Drug Interactions
Juan M. Banda, Tobias Kuhn, Nigam H. Shah, Michel Dumontier
ISWC (2)3
2015 Functional evaluation of out-of-the-box text-mining tools for data-mining tasks
abstract
OBJECTIVE: The trade-off between the speed and simplicity of dictionary-based term recognition and the richer linguistic information provided by more advanced natural language processing (NLP) is an area of active discussion in clinical informatics. In this paper, we quantify this trade-off among text processing systems that make different trade-offs between speed and linguistic understanding. We tested both types of systems in three clinical research tasks: phase IV safety profiling of a drug, learning adverse drug-drug interactions, and learning used-to-treat relationships between drugs and indications. MATERIALS: We first benchmarked the accuracy of the NCBO Annotator and REVEAL in a manually annotated, publically available dataset from the 2008 i2b2 Obesity Challenge. We then applied the NCBO Annotator and REVEAL to 9 million clinical notes from the Stanford Translational Research Integrated Database Environment (STRIDE) and used the resulting data for three research tasks. RESULTS: There is no significant difference between using the NCBO Annotator and REVEAL in the results of the three research tasks when using large datasets. In one subtask, REVEAL achieved higher sensitivity with smaller datasets. CONCLUSIONS: For a variety of tasks, employing simple term recognition methods instead of advanced NLP methods results in little or no impact on accuracy when using large datasets. Simpler dictionary-based methods have the advantage of scaling well to very large datasets. Promoting the use of simple, dictionary-based methods for population level analyses can advance adoption of NLP in practice.
Kenneth Jung, Paea LePendu, Srinivasan Iyer 0002, Anna Bauer-Mehren, Bethany Percha, Nigam H. Shah
J. Am. Medical Informatics Assoc.6
2015 A method for systematic discovery of adverse drug events from clinical notes
abstract
OBJECTIVE: Adverse drug events (ADEs) are undesired harmful effects resulting from use of a medication, and occur in 30% of hospitalized patients. The authors have developed a data-mining method for systematic, automated detection of ADEs from electronic medical records. MATERIALS AND METHODS: This method uses the text from 9.5 million clinical notes, along with prior knowledge of drug usages and known ADEs, as inputs. These inputs are further processed into statistics used by a discriminative classifier which outputs the probability that a given drug-disorder pair represents a valid ADE association. Putative ADEs identified by the classifier are further filtered for positive support in 2 independent, complementary data sources. The authors evaluate this method by assessing support for the predictions in other curated data sources, including a manually curated, time-indexed reference standard of label change events. RESULTS: This method uses a classifier that achieves an area under the curve of 0.94 on a held out test set. The classifier is used on 2,362,950 possible drug-disorder pairs comprised of 1602 unique drugs and 1475 unique disorders for which we had data, resulting in 240 high-confidence, well-supported drug-AE associations. Eighty-seven of them (36%) are supported in at least one of the resources that have information that was not available to the classifier. CONCLUSION: This method demonstrates the feasibility of systematic post-marketing surveillance for ADEs using electronic medical records, a key component of the learning healthcare system.
Kenneth Jung, Rainer Winnenburg, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2015 Implications of non-stationarity on predictive modeling using EHRs
Kenneth Jung, Nigam H. Shah
J. Biomed. Informatics2
2014 Medicine in the age of electronic health records
abstract
In the era of EHRs, it is possible to examine the outcomes of decisions made by doctors during clinical practice to identify patterns of care---generating evidence based on the collective practice of experts. We will discuss methods that use unstructured patient data to monitor for adverse drug events, profile specific drugs, identify off-label drug usage, uncover 'natural experiments' and generate practice-based evidence for difficult-to-test clinical hypotheses. We will describe how to detect associations among drugs and their adverse events several years before an alert is issued as well as compute the true rate of drug-drug interactions. We will present approaches to identify novel off-label uses of drugs using the patient feature matrix along with prior knowledge about drugs, diseases, and known usage. We will review a natural experiment--where a subset of congestive heart failure patients who were prescribed Cilostazol despite its black box warning--and profile its safety. We will discuss the testing of a clinical hypothesis about an association between allergic conditions and chronic uveitis in patients with juvenile idiopathic arthritis.
Nigam H. Shah
KDD1
2014 Finding progression stages in time-evolving event sequences
abstract
Event sequences, such as patients' medical histories or users' sequences of product reviews, trace how individuals progress over time. Identifying common patterns, or progression stages, in such event sequences is a challenging task because not every individual follows the same evolutionary pattern, stages may have very different lengths, and individuals may progress at different rates. In this paper, we develop a model-based method for discovering common progression stages in general event sequences. We develop a generative model in which each sequence belongs to a class, and sequences from a given class pass through a common set of stages, where each sequence evolves at its own rate. We then develop a scalable algorithm to infer classes of sequences, while also segmenting each sequence into a set of stages. We evaluate our method on event sequences, ranging from patients' medical histories to online news and navigational traces from the Web. The evaluation shows that our methodology can predict future events in a sequence, while also accurately inferring meaningful progression stages, and effectively grouping sequences based on common progression patterns. More generally, our methodology allows us to reason about how event sequences progress over time, by discovering patterns and categories of temporal evolution in large-scale datasets of events.
Jaewon Yang, Julian J. McAuley, Jure Leskovec, Paea LePendu, Nigam H. Shah
WWW5
2014 Toward personalizing treatment for depression: predicting diagnosis and severity
abstract
OBJECTIVE: Depression is a prevalent disorder difficult to diagnose and treat. In particular, depressed patients exhibit largely unpredictable responses to treatment. Toward the goal of personalizing treatment for depression, we develop and evaluate computational models that use electronic health record (EHR) data for predicting the diagnosis and severity of depression, and response to treatment. MATERIALS AND METHODS: We develop regression-based models for predicting depression, its severity, and response to treatment from EHR data, using structured diagnosis and medication codes as well as free-text clinical reports. We used two datasets: 35,000 patients (5000 depressed) from the Palo Alto Medical Foundation and 5651 patients treated for depression from the Group Health Research Institute. RESULTS: Our models are able to predict a future diagnosis of depression up to 12 months in advance (area under the receiver operating characteristic curve (AUC) 0.70-0.80). We can differentiate patients with severe baseline depression from those with minimal or mild baseline depression (AUC 0.72). Baseline depression severity was the strongest predictor of treatment response for medication and psychotherapy. CONCLUSIONS: It is possible to use EHR data to predict a diagnosis of depression up to 12 months in advance and to differentiate between extreme baseline levels of depression. The models use commonly available data on diagnosis, medication, and clinical progress notes, making them easily portable. The ability to automatically determine severity can facilitate assembly of large patient cohorts with similar severity from multiple sites, which may enable elucidation of the moderators of treatment response in the future.
Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah
J. Am. Medical Informatics Assoc.6
2014 Mining clinical text for signals of adverse drug-drug interactions
abstract
BACKGROUND AND OBJECTIVE: Electronic health records (EHRs) are increasingly being used to complement the FDA Adverse Event Reporting System (FAERS) and to enable active pharmacovigilance. Over 30% of all adverse drug reactions are caused by drug-drug interactions (DDIs) and result in significant morbidity every year, making their early identification vital. We present an approach for identifying DDI signals directly from the textual portion of EHRs. METHODS: We recognize mentions of drug and event concepts from over 50 million clinical notes from two sites to create a timeline of concept mentions for each patient. We then use adjusted disproportionality ratios to identify significant drug-drug-event associations among 1165 drugs and 14 adverse events. To validate our results, we evaluate our performance on a gold standard of 1698 DDIs curated from existing knowledge bases, as well as with signaling DDI associations directly from FAERS using established methods. RESULTS: Our method achieves good performance, as measured by our gold standard (area under the receiver operator characteristic (ROC) curve >80%), on two independent EHR datasets and the performance is comparable to that of signaling DDIs from FAERS. We demonstrate the utility of our method for early detection of DDIs and for identifying alternatives for risky drug combinations. Finally, we publish a first of its kind database of population event rates among patients on drug combinations based on an EHR corpus. CONCLUSIONS: It is feasible to identify DDI signals and estimate the rate of adverse events among patients on drug combinations, directly from clinical text; this could have utility in prioritizing drug interaction surveillance as well as in clinical decision support.
Srinivasan Iyer 0002, Rave Harpaz, Paea LePendu, Anna Bauer-Mehren, Nigam H. Shah
J. Am. Medical Informatics Assoc.5
2013 Predictive Models in Mental Health: From Diagnosis to Treatment
Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah
AMIA6
2013 Learning Practice-based Evidence from Unstructured Clinical Notes
Nigam H. Shah, Paea LePendu, Anna Bauer-Mehren, Srinivasan Iyer 0002, Kenneth Jung, Tyler Cole, Rave Harpaz
AMIA1
2013 Mining Biomedical Ontologies and Data Using RDF Hypergraphs
abstract
As researchers analyze huge amounts of data that are annotated by large biomedical ontologies, one of the major challenges for data mining and machine learning is to leverage both ontologies and data together in a systematic and scalable way. In this paper, we address two interesting and related problems for mining biomedical ontologies and data: i) how to discover semantic associations with the help of formal ontologies, ii) how to identify potential errors in the ontologies with the help of data. By representing both ontologies and data using RDF hyper graphs, and subsequently transforming the hyper graphs to corresponding bipartite forms, we provide a generalized data mining method that scales beyond what existing ontology-based approaches can provide. We show the proposed method is indeed capable of capturing semantic associations while seamlessly incorporate domain knowledge in ontologies by performing evaluations on real-world electronic health dataset and NCBO ontologies. We also show that our data mining methods can discover and suggest corrections for misinformation in biomedical ontologies.
Haishan Liu, Dejing Dou, Ruoming Jin, Paea LePendu, Nigam H. Shah
ICMLA (1)5
2013 Empirical bayes model to combine signals of adverse drug reactions
abstract
Data mining is a crucial tool for identifying risk signals of potential adverse drug reactions (ADRs). However, mining of ADR signals is currently limited to leveraging a single data source at a time. It is widely believed that combining ADR evidence from multiple data sources will result in a more accurate risk identification system. We present a methodology based on empirical Bayes modeling to combine ADR signals mined from ~5 million adverse event reports collected by the FDA, and healthcare data corresponding to 46 million patients' the main two types of information sources currently employed for signal detection. Based on four sets of test cases (gold standard), we demonstrate that our method leads to a statistically significant and substantial improvement in signal detection accuracy, averaging 40% over the use of each source independently, and an area under the ROC curve of 0.87. We also compare the method with alternative supervised learning approaches, and argue that our approach is preferable as it does not require labeled (training) samples whose availability is currently limited. To our knowledge, this is the first effort to combine signals from these two complementary data sources, and to demonstrate the benefits of a computationally integrative strategy for drug safety surveillance.
Rave Harpaz, William DuMouchel, Paea LePendu, Nigam H. Shah
KDD4
2013 STOP using just GO: a multi-ontology hypothesis generation tool for high throughput experimentation
abstract
BACKGROUND: Gene Ontology (GO) enrichment analysis remains one of the most common methods for hypothesis generation from high throughput datasets. However, we believe that researchers strive to test other hypotheses that fall outside of GO. Here, we developed and evaluated a tool for hypothesis generation from gene or protein lists using ontological concepts present in manually curated text that describes those genes and proteins. RESULTS: As a consequence we have developed the method Statistical Tracking of Ontological Phrases (STOP) that expands the realm of testable hypotheses in gene set enrichment analyses by integrating automated annotations of genes to terms from over 200 biomedical ontologies. While not as precise as manually curated terms, we find that the additional enriched concepts have value when coupled with traditional enrichment analyses using curated terms. CONCLUSION: Multiple ontologies have been developed for gene and protein annotation, by using a dataset of both manually curated GO terms and automatically recognized concepts from curated text we can expand the realm of hypotheses that can be discovered. The web application STOP is available at http://mooneygroup.org/stop/.
Tobias Wittkop, Emily TerAvest, Uday S. Evani, K. Mathew Fleisch, Ari E. Berman, Corey Powell, Nigam H. Shah, Sean D. Mooney
BMC Bioinform.7
2013 Combing signals from spontaneous reports and electronic health records for detection of adverse drug reactions
abstract
OBJECTIVE: Data-mining algorithms that can produce accurate signals of potentially novel adverse drug reactions (ADRs) are a central component of pharmacovigilance. We propose a signal-detection strategy that combines the adverse event reporting system (AERS) of the Food and Drug Administration and electronic health records (EHRs) by requiring signaling in both sources. We claim that this approach leads to improved accuracy of signal detection when the goal is to produce a highly selective ranked set of candidate ADRs. MATERIALS AND METHODS: Our investigation was based on over 4 million AERS reports and information extracted from 1.2 million EHR narratives. Well-established methodologies were used to generate signals from each source. The study focused on ADRs related to three high-profile serious adverse reactions. A reference standard of over 600 established and plausible ADRs was created and used to evaluate the proposed approach against a comparator. RESULTS: The combined signaling system achieved a statistically significant large improvement over AERS (baseline) in the precision of top ranked signals. The average improvement ranged from 31% to almost threefold for different evaluation categories. Using this system, we identified a new association between the agent, rasburicase, and the adverse event, acute pancreatitis, which was supported by clinical review. CONCLUSIONS: The results provide promising initial evidence that combining AERS with EHRs via the framework of replicated signaling can improve the accuracy of signal detection for certain operating scenarios. The use of additional EHR data is required to further evaluate the capacity and limits of this system and to extend the generalizability of these results.
Rave Harpaz, Santiago Vilar, William DuMouchel, Hojjat Salmasian, Krystl Haerian, Nigam H. Shah, Herbert S. Chase, Carol Friedman
J. Am. Medical Informatics Assoc.6
2013 Web-scale pharmacovigilance: listening to signals from the crowd
abstract
Adverse drug events cause substantial morbidity and mortality and are often discovered after a drug comes to market. We hypothesized that Internet users may provide early clues about adverse drug events via their online information-seeking. We conducted a large-scale study of Web search log data gathered during 2010. We pay particular attention to the specific drug pairing of paroxetine and pravastatin, whose interaction was reported to cause hyperglycemia after the time period of the online logs used in the analysis. We also examine sets of drug pairs known to be associated with hyperglycemia and those not associated with hyperglycemia. We find that anonymized signals on drug interactions can be mined from search logs. Compared to analyses of other sources such as electronic health records (EHR), logs are inexpensive to collect and mine. The results demonstrate that logs of the search activities of populations of computer users can contribute to drug safety surveillance.
Ryen W. White, Nicholas P. Tatonetti, Nigam H. Shah, Russ B. Altman, Eric Horvitz
J. Am. Medical Informatics Assoc.3
2012 Mining the pharmacogenomics literature - a survey of the state of the art
abstract
This article surveys efforts on text mining of the pharmacogenomics literature, mainly from the period 2008 to 2011. Pharmacogenomics (or pharmacogenetics) is the field that studies how human genetic variation impacts drug response. Therefore, publications span the intersection of research in genotypes, phenotypes and pharmacology, a topic that has increasingly become a focus of active research in recent years. This survey covers efforts dealing with the automatic recognition of relevant named entities (e.g. genes, gene variants and proteins, diseases and other pathological phenomena, drugs and other chemicals relevant for medical treatment), as well as various forms of relations between them. A wide range of text genres is considered, such as scientific publications (abstracts, as well as full texts), patent texts and clinical narratives. We also discuss infrastructure and resources needed for advanced text analytics, e.g. document corpora annotated with corresponding semantic metadata (gold standards and training data), biomedical terminologies and ontologies providing domain-specific background knowledge at different levels of formality and specificity, software architectures for building complex and scalable text analytics pipelines and Web services grounded to them, as well as comprehensive ways to disseminate and interact with the typically huge amounts of semiformal knowledge structures extracted by text mining tools. Finally, we consider some of the novel applications that have already been developed in the field of pharmacogenomic text mining and point out perspectives for future research.
Udo Hahn, Kevin Cohen 0001, Yael Garten, Nigam H. Shah
Briefings Bioinform.4
2012 Using ontology-based annotation to profile disease research
abstract
BACKGROUND: Profiling the allocation and trend of research activity is of interest to funding agencies, administrators, and researchers. However, the lack of a common classification system hinders the comprehensive and systematic profiling of research activities. This study introduces ontology-based annotation as a method to overcome this difficulty. Analyzing over a decade of funding data and publication data, the trends of disease research are profiled across topics, across institutions, and over time. RESULTS: This study introduces and explores the notions of research sponsorship and allocation and shows that leaders of research activity can be identified within specific disease areas of interest, such as those with high mortality or high sponsorship. The funding profiles of disease topics readily cluster themselves in agreement with the ontology hierarchy and closely mirror the funding agency priorities. Finally, four temporal trends are identified among research topics. CONCLUSIONS: This work utilizes disease ontology (DO)-based annotation to profile effectively the landscape of biomedical research activity. By using DO in this manner a use-case driven mechanism is also proposed to evaluate the utility of classification hierarchies.
Adrien Coulet, Paea LePendu, Nigam H. Shah
J. Am. Medical Informatics Assoc.4
2012 The National Center for Biomedical Ontology
abstract
The National Center for Biomedical Ontology is now in its seventh year. The goals of this National Center for Biomedical Computing are to: create and maintain a repository of biomedical ontologies and terminologies; build tools and web services to enable the use of ontologies and terminologies in clinical and translational research; educate their trainees and the scientific community broadly about biomedical ontology and ontology-based technology and best practices; and collaborate with a variety of groups who develop and use ontologies and terminologies in biomedicine. The centerpiece of the National Center for Biomedical Ontology is a web-based resource known as BioPortal. BioPortal makes available for research in computationally useful forms more than 270 of the world's biomedical ontologies and terminologies, and supports a wide range of web services that enable investigators to use the ontologies to annotate and retrieve data, to generate value sets and special-purpose lexicons, and to perform advanced analytics on a wide range of biomedical data.
Mark A. Musen, Natasha F. Noy, Nigam H. Shah, Patricia L. Whetzel, Christopher G. Chute, Margaret-Anne D. Storey, Barry Smith 0001
J. Am. Medical Informatics Assoc.3
2012 The coming age of data-driven medicine: translational bioinformatics' next frontier
abstract
Last year, in 2011, we argued that biomedical informatics stands ready to revolutionize human health and healthcare using large-scale measurements on a large number of individuals.1 We anticipated that, with the coming changes in the amount and diversity of datasets, data-centric approaches that compute on massive amounts of data (often called ‘Big Data’2,3) to discover patterns and to make clinically relevant predictions would be increasingly common in translational bioinformatics. Given these trends, we programmed the 2012 Summit on Translational Bioinformatics to focus on research that takes us from base pairs to the bedside,4 with a particular emphasis on clinical implications of mining massive datasets, and bridging the latest multimodal measurement technologies with the large amounts of electronic healthcare data that are increasingly available. The coming year did turn out to be the year of Big Data for the Summit, with multiple submissions on managing and interpreting large datasets (figure 1). Among the 35 full paper submissions to the Summit, four stood out for their innovation, and hence the authors were invited to expand the work for this special issue of JAMIA—adding to the growing presence of translational bioinformatics in the journal.5–9 A tag cloud generated from the title and abstracts of the submissions made to the AMIA Translational Bioinformatics Summit 2012. The more frequently used the words are, the larger they appear. ‘Data’ was the most commonly mentioned word across all submissions for 2012. Liu et al10 demonstrated how the ability to predict adverse drug reactions can be increased by integrating chemical, biological, and phenotypic properties of drugs. They demonstrated that prediction accuracy increased from 0.9054 (when only chemical structures were used) to 0.9524 (when chemical structures along with biological and phenotypic features were used). They conclude that data fusion approaches are promising for large-scale adverse drug reaction predictions in both preclinical and post-marketing phases. Bhavnani et al11 assert that existing methods to analyze ancestral informative single-nucleotide polymorphisms (SNPs) (ie, SNPs that have large differences in genotype frequencies between two or more ancestral populations) identify a parsimonious set of SNPs that can identify distinct population clusters. However, existing methods do not directly visualize which clusters of subjects are related to which clusters of SNPs, or allow visualization of the genotypes that determine the cluster memberships. In an attempt to reveal such hidden relationships, they used three bipartite analytical representations (a bipartite network, a heat map with dendrograms, and a Circos ideogram) to simultaneously visualize clusters of subjects, SNPs, and the attributes that cause them to cluster. Seeking to maximize the utility of the abundance of available genome-wide association study (GWAS) data, Russu et al12 introduced a novel Bayesian model search algorithm, binary outcome stochastic search, for model selection when the number of predictors (eg, SNPs) far exceeds the number of observations. They propose an innovative stochastic model search technique where the relationship between the observed responses and the available predictors is described by a latent variable model with a probit link. They compare binary outcome stochastic search with three established methods (stepwise regression, logistic lasso, and elastic net) in a simulated study and in two real world studies to demonstrate higher precision (while preserving recall) in identifying SNPs associated with the observed outcome than the one obtained from established methods. Morgan et al,13 recipient of the Marco Ramoni Best Paper Award, constructed genomic disease risk summaries for 55 common diseases using reported gene–disease associations in the research literature. They constructed risk profiles based on the SNPs as well as on 187 whole-genome sequences and show that risk predictions derived from sequencing differ substantially from those obtained from the SNPs for several different non-monogenic diseases. When a large fraction of associated variants for a given disease is not covered by the genotyping array, the overall risk predictions can vary dramatically—by as much as a factor of 20 times in some instances. Beyond this year's conference papers, in the larger informatics community, researchers have demonstrated that GWAS can now be performed by leveraging large amounts of electronic medical record (EMR) data. For example, Kho et al showed that, by using commonly available data from five different EMRs, it is possible to accurately identify type 2 diabetes cases and controls for genetic study across multiple institutions.14 In addition, genomic sequencing has moved out of the research realm and established itself in the clinic. For example, at the Medical College of Wisconsin, Dr Howard Jacob's team used genome sequencing to identify a novel causal mutation that led to successful treatment of a 6-year-old boy with an extreme form of inflammatory bowel disease.15,16 Currently, the discussion of Big Data in translational informatics often connotes next-generation sequencing data.3,17,18 However, this is beginning to change: in 2011, the use of large public datasets of various kinds increased dramatically. The research activity around data mining for predicting adverse drug events (ADEs) using public data is an excellent example.19 Drug safety surveillance is currently based on spontaneous reporting systems, which contain reports of suspected ADEs seen in clinical practice. In the USA, the primary database for such reports is the Adverse Event Reporting System (AERS) database at the Food and Drug Administration. This resource has been successfully mined using ‘disproportionality measures’, which quantify the magnitude of difference between observed and expected rates of particular drug–ADE pairs.20,21 Given the amount of data available in AERS,22 researchers are developing methods for detecting new or latent multi-drug adverse events. Examples include using side effect profiles from AERs' reports to infer the presence of unreported adverse events,23–25 and creating a network of known drug–ADE relationships to predict as yet unknown ADEs before they are found in post-market evidence.26 Going beyond reported adverse events and making use of molecular level data, Pouliot et al27 generated logistic regression models to correlate and predict post-marketing ADEs based on screening data from PubChem, a public database of chemical structures of small organic molecules along with information about their biological activities. In a related effort, Vilar et al28 devised a way to enhance existing, data-mining algorithms with chemical information using molecular fingerprints—which represent molecules through a bit vector that codifies the existence of particular structural features or functional groups—to enhance ADE signals generated from adverse event reports. There have been increasing efforts to use other data sources, such as EMRs, for the purpose of detecting ADEs29–31 and to discover multi-drug ADEs.32 Researchers have also used billing and claims data for active drug safety surveillance33–35 and applied literature mining for drug safety.36 Recently, Chee et al37 explored the use of online health forums as a source of data to identify drugs for further scrutiny. They aggregate individuals' opinions of drugs in roughly 12 million personal health messages using natural language processing and are able to identify drugs withdrawn from the market based on messages discussing them before their removal. Looking ahead, we believe that Big Data in biomedical informatics will be far more than genome sequence data.38–40 We argue that Big Data must be considered in a comprehensive manner, including both large amounts of ‘molecular measurements' on a person (eg, sequencing) and small amounts of ‘routine measurements' on a large number of people (eg, clinical notes, laboratory measurements, claims data and adverse event reports). In contrast with the buzz around genomic-data-in-the-clinic or adverse event predictions, consider the example by Frankovich et al.41 When the existing literature and a survey of colleagues was insufficient to guide the clinical care of a patient, Frankovich et al applied trend analysis to the EMR data from 98 patients to ‘learn’ a data-driven guideline on how to provide care for a 13-year-old girl with systemic lupus erythematosus.41 Such data-centric approaches are particularly useful when derivation of a formal guideline is not feasible from a practical standpoint. It is tantalizing to imagine how scientific inquiry would be performed differently if we collect and share access to lots of data—both genomic and ‘routine’. How will the kinds of questions we ask change when we cross a certain data threshold?42,43 For example, researchers at Carnegie Mellon University built a scene completion tool by scraping millions of other images on the web from public sources. After the system accumulated a corpus of millions of photos, completed scenes were indistinguishable to the naked eye. The case for Big Data analytics has already won over the legal domain in at least one application, replacing armies of lawyers with computer algorithms designed for ‘e-discovery’—that is, retrieval of relevant materials for a legal case.44 Even the liberal arts are embracing Big Data: capitalizing on Google's efforts to digitize books, researchers in the humanities are blazing new trails in ‘culturomics' by examining language based on the analysis of word combinations occurring in millions of digitized books through time.45 In 2013, we will have the sixth Summit on Translational Bioinformatics and the third year of the AMIA Joint Summits on Translational Science. Translational research has become integral to the biomedical research enterprise, as evidenced by the creation of a National Center for Advancing Translational Science at the NIH. The Joint Summits continue to be a venue to facilitate dramatic changes that are underway to deliver quality, personalized healthcare in the USA without increasing spending at a rate exceeding the growth of the GDP.46 Reflecting this priority, the 2013 TBI Summit will have new tracks that will showcase the ways in which the translational sciences are having a significant impact on the way clinical care, biomedical research, and drug discovery are performed. We believe that the time is ripe for medicine to embrace Big Data, to usher in the age of data-driven medicine—and to truly enable proactive, predictive, preventive, participatory, and patient-centered health.47 Data-driven medicine will enable the discovery of new treatment options based on the multi-model molecular measurements on patients and learning from the trends hidden among the diagnoses, prescriptions, and discharge summaries of millions of patient encounters logged by clinical practitioners.48,49 The increasing synergy between the Translational Bioinformatics Summit and the Clinical Research Informatics Summit is an indication of this impending convergence. This is an exciting time when medicine begins utilizing massive amounts of data to discover patterns and trends and to make predictions in a manner that is a mainstay of web-scale computing.42 NHS is funded by the US National Institute of Health Roadmap (U54 HG004028 and U54 LM008748). JDT is funded by a Clinical and Translational Science Award (UL1 RR024128) and a gift from David H Murdock. None. Commissioned; internally peer reviewed.
Nigam H. Shah, Jessica D. Tenenbaum
J. Am. Medical Informatics Assoc.1
2012 Unified Medical Language System term occurrences in clinical notes: a large-scale corpus analysis
abstract
OBJECTIVE: To characterise empirical instances of Unified Medical Language System (UMLS) Metathesaurus term strings in a large clinical corpus, and to illustrate what types of term characteristics are generalisable across data sources. DESIGN: Based on the occurrences of UMLS terms in a 51 million document corpus of Mayo Clinic clinical notes, this study computes statistics about the terms' string attributes, source terminologies, semantic types and syntactic categories. Term occurrences in 2010 i2b2/VA text were also mapped; eight example filters were designed from the Mayo-based statistics and applied to i2b2/VA data. RESULTS: For the corpus analysis, negligible numbers of mapped terms in the Mayo corpus had over six words or 55 characters. Of source terminologies in the UMLS, the Consumer Health Vocabulary and Systematized Nomenclature of Medicine-Clinical Terms (SNOMED-CT) had the best coverage in Mayo clinical notes at 106426 and 94788 unique terms, respectively. Of 15 semantic groups in the UMLS, seven groups accounted for 92.08% of term occurrences in Mayo data. Syntactically, over 90% of matched terms were in noun phrases. For the cross-institutional analysis, using five example filters on i2b2/VA data reduces the actual lexicon to 19.13% of the size of the UMLS and only sees a 2% reduction in matched terms. CONCLUSION: The corpus statistics presented here are instructive for building lexicons from the UMLS. Features intrinsic to Metathesaurus terms (well formedness, length and language) generalise easily across clinical institutions, but term frequencies should be adapted with caution. The semantic groups of mapped terms may differ slightly from institution to institution, but they differ greatly when moving to the biomedical literature domain.
Stephen T. Wu, Dingcheng Li, Cui Tao, Mark A. Musen, Christopher G. Chute, Nigam H. Shah
J. Am. Medical Informatics Assoc.7
2012 Chapter 9: Analyses Using Disease Ontologies
abstract
Advanced statistical methods used to analyze high-throughput data such as gene-expression assays result in long lists of "significant genes." One way to gain insight into the significance of altered expression levels is to determine whether Gene Ontology (GO) terms associated with a particular biological process, molecular function, or cellular component are over- or under-represented in the set of genes deemed significant. This process, referred to as enrichment analysis, profiles a gene-set, and is widely used to makes sense of the results of high-throughput experiments. The canonical example of enrichment analysis is when the output dataset is a list of genes differentially expressed in some condition. To determine the biological relevance of a lengthy gene list, the usual solution is to perform enrichment analysis with the GO. We can aggregate the annotating GO concepts for each gene in this list, and arrive at a profile of the biological processes or mechanisms affected by the condition under study. While GO has been the principal target for enrichment analysis, the methods of enrichment analysis are generalizable. We can conduct the same sort of profiling along other ontologies of interest. Just as scientists can ask "Which biological process is over-represented in my set of interesting genes or proteins?" we can also ask "Which disease (or class of diseases) is over-represented in my set of interesting genes or proteins?". For example, by annotating known protein mutations with disease terms from the ontologies in BioPortal, Mort et al. recently identified a class of diseases--blood coagulation disorders--that were associated with a 14-fold depletion in substitutions at O-linked glycosylation sites. With the availability of tools for automatic annotation of datasets with terms from disease ontologies, there is no reason to restrict enrichment analyses to the GO. In this chapter, we will discuss methods to perform enrichment analysis using any ontology available in the biomedical domain. We will review the general methodology of enrichment analysis, the associated challenges, and discuss the novel translational analyses enabled by the existence of public, national computational infrastructure and by the use of disease ontologies in such analyses.
Nigam H. Shah, Tyler Cole, Mark A. Musen
PLoS Comput. Biol.1
2011 Computationally translating molecular discoveries into tools for medicine: translational bioinformatics articles now featured in JAMIA
abstract
This year marks the 15th anniversary of the invention of the gene expression microarray. As mRNA transcripts serve as the blueprint within cells for making proteins, measuring mRNA levels was seen as an accurate and manageable way to investigate cell and tissue processes. Those earliest microarrays in 1995 could measure 48 transcripts in parallel in plants,1 but within 1 year were scaled up to measure more than 1000 transcripts including those in human tissues. Today, these microarrays are essentially commodity items, commonly used to study human health and disease in hospitals and academic institutions, as well as in the biotechnology and pharmaceutical industry. While tens of thousands of publications have already been published referencing microarrays, this is just the start. Similar arrays are already used to probe genetic differences in DNA, but even these will soon be supplanted by whole genome sequencing, where we can expect all three billion human base pairs to be sequenced for a few thousand dollars. The exponential decrease in costs for whole genome sequencing has been described as going beyond the decline we are used to from Moore's law.2 These enormous amounts of molecular data make it clear that there is a pressing need for computational methods to analyze and interpret them. Molecular data have never been foreign to the pages of JAMIA or AMIA Symposia. The second volume of JAMIA back in 1995 contained an article describing how the internet and newly introduced world wide web could be used to facilitate genome sequencing efforts across two academic genome centers and introduced concepts like yeast artificial chromosomes, sequence tagged sites and contigs.3 The 2002 AMIA Fall Symposium suggested bioinformatics and medical informatics could and should be part of the same discipline of biomedical informatics. While not every publication in bioinformatics is likely relevant to AMIA members, it is important to track those applying bioinformatics to human health and disease. Translational bioinformatics can be defined as ‘the development of storage, analytic, and interpretive methods to optimize the transformation of increasingly voluminous biomedical data into proactive, predictive, preventive, and participatory health.’4 Indeed, translational bioinformatics has been a core strategic area of importance for AMIA since 2008. Today, as more hospital and academic medical centers embrace molecular measurements for diagnosis and planning therapies, we recognize that the community of JAMIA readers will need to keep abreast of new developments and applications of translational bioinformatics.5 In the past few years, however, it has been increasingly hard to find articles on translational bioinformatics in JAMIA. However, with this month's issue, we hope to start reversing the trend. Starting with a Perspective on from Neil Sarkar and colleagues (see page 354), we are highlighting five manuscripts in the field of translational bioinformatics, and ‘opening the doors' for submissions from investigators and authors in this field.6 Recognizing the continued growth in bioinformatics, especially as related to human health and disease, in 2009 AMIA initiated the annual Summit on Translational Bioinformatics as a new annual meeting to address the growing need for a scientific conference to present and discuss developments and the application of methods in this field. At the 2011 Summit, several authors of top-ranked submitted poster abstracts were invited to expand their submissions into full manuscripts, which were then evaluated by JAMIA reviewers. One such manuscript appears in this month's issue. Hua Xu and colleagues (see page 387) show how dosing details locked within free-text electronic health records could be found using natural language processing, thus enabling a gene–dosing association study for warfarin dosing at Vanderbilt University.7 This work serves as a premier model of the type of translational research possible when bioinformatics and medical informatics investigators truly collaborate. Few institutions have the considerable resources of Vanderbilt University, which has a DNA biobank linked to a de-identified electronic health record subset, but there are a few, notably the Mayo Clinic. In a manuscript this month, Pathak and colleagues (see page 376) show how standardized representations of clinical characteristics extracted from institutional electronic health records can be integrated to enable large scale genetics studies between Vanderbilt University and the Mayo Clinic.8 As the genetic risk of disease is now theorized to result from a combination of rare DNA variants,9 larger cross-country cohorts like these will be needed to identify these variants. Integration of information on patients, samples, or data is indeed a common theme in translational bioinformatics. David Foran and colleagues (see page 403) highlight a new federated software system that can handle another type of biobank for pathology samples.10 Processed into histology cores and distributed on a tissue microarray, these samples are proving to be invaluable for the high-throughput query of specific proteins. James Chen and colleagues (see page 392) show how comparing the genomic differences found across publicly available data from previous prostate cancer studies can yield a single core ‘signature’ that can distinguish prostate cancer patients with better prognosis from those with worse prognosis.11 Translational bioinformatics grew out of the work of a small but cohesive group of researchers who bridged the gap between computational biology and medicine. It is with sad remembrance that we note the passing of Marco Ramoni, a pioneering spirit who was one of the first Track Chairs for the AMIA Summit on Translational Bioinformatics. So that his contributions to AMIA and the field of translational bioinformatics will not be forgotten, this year the Board of Directors established the Marco Ramoni Best Paper Award, to be presented each year at the AMIA Summit on Translational Bioinformatics. The manuscript from the award winner for 2011, Wei Wei, appears in this issue of JAMIA (see page 370).12 Selected by an external review panel chosen by the Chair of the Scientific Program Committee and again peer reviewed by JAMIA reviewers, it is slightly ironic that the award winning manuscript is on the application of Bayes' theorem, coincidentally similar to Dr Ramoni's own work, which is summarized by Kohane and Szolovits (see page 367) in an invited academic tribute to our dear colleague.13 Next year will see the fifth Summit on Translational Bioinformatics. With the approaching changes in the amount and diversity of datasets discussed above, we anticipate that data-centric approaches that compute on massive amounts of data (often called ‘big data’14) to identify patterns and make clinically relevant predictions will be increasingly common in translational bioinformatics. In anticipation, the 2012 Summit on Translational Bioinformatics will have four tracks focusing on research that take us from base pairs to the bedside,15 with a particular emphasis on the clinical implications of mining massive datasets, and bridging the latest multimodal measurement technologies using the large amounts of electronic healthcare data that are increasingly available. We invite readers to submit extended (10-page) submissions for the Summit before the deadline of August 15. The top papers will be published in JAMIA after peer review and the editorial office will make all efforts to have them available online first by the time the Summit takes place on March 19–21, 2012 in San Francisco. In closing, we have continued to note arguments over how much translational bioinformatics informatics investigators and professionals really need to know. To answer this, we must consider that we are entering a decade where tens of thousands of people have already obtained samplings of their own DNA sequences from consumer genomics companies,16 with over 700 000 RNA microarray measurements already available to the public,17,18 and we expect that 30 000 people will have their whole genome sequenced this year alone.19 We must not keep assuming that if we just build and provide the right generalized tools and methods, others will take them and use them the right way to improve healthcare—the field is over-saturated with tools already. If we best understand the tools and methods we build, we need to be the first to actually use those tools, and show the world what can be achieved. Instead of discussing how relevant translational bioinformatics is, we need to argue that biomedical informatics is the only field in biomedicine that is ready to revolutionize human health and healthcare using these tools and measurements. AJB is funded by the US National Library of Medicine (R01 LM009719) and the Lucile Packard Foundation for Children's Health. NHS is funded by the US National Institute of Health Roadmap (U54 HG004028). None. Commissioned; internally peer reviewed.
Atul J. Butte, Nigam H. Shah
J. Am. Medical Informatics Assoc.2
2011 Enabling enrichment analysis with the Human Disease Ontology
Paea LePendu, Mark A. Musen, Nigam H. Shah
J. Biomed. Informatics3
2011 NCBO Resource Index: Ontology-based search and mining of biomedical resources
Clément Jonquet, Paea LePendu, Sean M. Falconer, Adrien Coulet, Natasha F. Noy, Mark A. Musen, Nigam H. Shah
J. Web Semant.7
2010 Optimize First, Buy Later: Analyzing Metrics to Ramp-Up Very Large Knowledge Bases
Paea LePendu, Natasha F. Noy, Clément Jonquet, Paul R. Alexander, Nigam H. Shah, Mark A. Musen
ISWC (1)5
2010 A UIMA wrapper for the NCBO annotator
abstract
SUMMARY: The Unstructured Information Management Architecture (UIMA) framework and web services are emerging as useful tools for integrating biomedical text mining tools. This note describes our work, which wraps the National Center for Biomedical Ontology (NCBO) Annotator-an ontology-based annotation service-to make it available as a component in UIMA workflows. AVAILABILITY: This wrapper is freely available on the web at http://bionlp-uima.sourceforge.net/ as part of the UIMA tools distribution from the Center for Computational Pharmacology (CCP) at the University of Colorado School of Medicine. It has been implemented in Java for support on Mac OS X, Linux and MS Windows.
Christophe Roeder, Clément Jonquet, Nigam H. Shah, William A. Baumgartner Jr., Karin Verspoor, Lawrence Hunter
Bioinform.3
2010 Using text to build semantic networks for pharmacogenomics
Adrien Coulet, Nigam H. Shah, Yael Garten, Mark A. Musen, Russ B. Altman
J. Biomed. Informatics2
2009 What Four Million Mappings Can Tell You about Two Hundred Ontologies
Amir Ghazvinian, Natasha F. Noy, Clément Jonquet, Nigam H. Shah, Mark A. Musen
ISWC4
2009 Comparison of concept recognizers for building the Open Biomedical Annotator
abstract
The National Center for Biomedical Ontology (NCBO) is developing a system for automated, ontology-based access to online biomedical resources (Shah NH, et al.: Ontology-driven indexing of public datasets for translational bioinformatics. BMC Bioinformatics 2009, 10(Suppl 2):S1). The system's indexing workflow processes the text metadata of diverse resources such as datasets from GEO and ArrayExpress to annotate and index them with concepts from appropriate ontologies. This indexing requires the use of a concept-recognition tool to identify ontology concepts in the resource's textual metadata. In this paper, we present a comparison of two concept recognizers - NLM's MetaMap and the University of Michigan's Mgrep. We utilize a number of data sources and dictionaries to evaluate the concept recognizers in terms of precision, recall, speed of execution, scalability and customizability. Our evaluations demonstrate that Mgrep has a clear edge over MetaMap for large-scale service oriented applications. Based on our analysis we also suggest areas of potential improvements for Mgrep. We have subsequently used Mgrep to build the Open Biomedical Annotator service. The Annotator service has access to a large dictionary of biomedical terms derived from the United Medical Language System (UMLS) and NCBO ontologies. The Annotator also leverages the hierarchical structure of the ontologies and their mappings to expand annotations. The Annotator service is available to the community as a REST Web service for creating ontology-based annotations of their data.
Nigam H. Shah, Nipun Bhatia, Clément Jonquet, Daniel L. Rubin, Annie P. Chiang, Mark A. Musen
BMC Bioinform.1
2009 Ontology-driven indexing of public datasets for translational bioinformatics
abstract
The volume of publicly available genomic scale data is increasing. Genomic datasets in public repositories are annotated with free-text fields describing the pathological state of the studied sample. These annotations are not mapped to concepts in any ontology, making it difficult to integrate these datasets across repositories. We have previously developed methods to map text-annotations of tissue microarrays to concepts in the NCI thesaurus and SNOMED-CT. In this work we generalize our methods to map text annotations of gene expression datasets to concepts in the UMLS. We demonstrate the utility of our methods by processing annotations of datasets in the Gene Expression Omnibus. We demonstrate that we enable ontology-based querying and integration of tissue and gene expression microarray data. We enable identification of datasets on specific diseases across both repositories. Our approach provides the basis for ontology-driven data integration for translational research on gene and protein expression data. Based on this work we have built a prototype system for ontology based annotation and indexing of biomedical data. The system processes the text metadata of diverse resource elements such as gene expression data sets, descriptions of radiology images, clinical-trial reports, and PubMed article abstracts to annotate and index them with concepts from appropriate ontologies. The key functionality of this system is to enable users to locate biomedical data resources related to particular ontology concepts.
Nigam H. Shah, Clément Jonquet, Annie P. Chiang, Atul J. Butte, Rong Chen 0006, Mark A. Musen
BMC Bioinform.1
2008 Comparison of Ontology-based Semantic-Similarity Measures
Wei-Nchih Lee, Nigam H. Shah, Karanjot Sundlass, Mark A. Musen
AMIA2
2008 UMLS-Query: A Perl Module for Querying the UMLS
Nigam H. Shah, Mark A. Musen
AMIA1
2008 Biomedical ontologies: a functional perspective
abstract
The information explosion in biology makes it difficult for researchers to stay abreast of current biomedical knowledge and to make sense of the massive amounts of online information. Ontologies--specifications of the entities, their attributes and relationships among the entities in a domain of discourse--are increasingly enabling biomedical researchers to accomplish these tasks. In fact, bio-ontologies are beginning to proliferate in step with accruing biological data. The myriad of ontologies being created enables researchers not only to solve some of the problems in handling the data explosion but also introduces new challenges. One of the key difficulties in realizing the full potential of ontologies in biomedical research is the isolation of various communities involved: some workers spend their career developing ontologies and ontology-related tools, while few researchers (biologists and physicians) know how ontologies can accelerate their research. The objective of this review is to give an overview of biomedical ontology in practical terms by providing a functional perspective--describing how bio-ontologies can and are being used. As biomedical scientists begin to recognize the many different ways ontologies enable biomedical research, they will drive the emergence of new computer applications that will help them exploit the wealth of research data now at their fingertips.
Daniel L. Rubin, Nigam H. Shah, Natasha F. Noy
Briefings Bioinform.2
2007 Interpretation Errors related to the GO Annotation File Format
Dilvan de Abreu Moreira, Nigam H. Shah, Mark A. Musen
AMIA2
2007 Searching ontologies based on content: experiments in the biomedical domain
abstract
As more ontologies become publicly available, finding the "right" ontologies becomes much harder. In this paper, we address the problem of ontology search: finding a collection of ontologies from an ontology repository that are relevant to the user's query. In particular, we look at the case when users search for ontologies relevant to a particular topic (e.g., an ontology about anatomy). Ontologies that are most relevant to such query often do not have the query term in the names of their concepts (e.g., the Foundational Model of Anatomy ontology does not have the term "anatomy" in any of its concepts' names). Thus, we present a new ontology-search technique that helps users in these types of searches. When looking for ontologies on a particular topic (e.g., anatomy), we retrieve from the Web a collection of terms that represent the given domain (e.g., terms such as body, brain, skin, etc. for anatomy). We then use these terms to expand the user query. We evaluate our algorithm on queries for topics in the biomedical domain against a repository of biomedical ontologies. We use the results obtained from experts in the biomedical-ontology domain as the gold standard. Our experiments demonstrate that using our method for query expansion improves retrieval results by a 113%, compared to the tools that search only for the user query terms and consider only class and property names (like Swoogle). We show 43% improvement for the case where not only class and property names but also property values are taken into account.
Harith Alani, Natasha F. Noy, Nigam H. Shah, Nigel Shadbolt, Mark A. Musen
K-CAP3
2007 Current progress in network research: toward reference networks for key model organisms
abstract
The collection of multiple genome-scale datasets is now routine, and the frontier of research in systems biology has shifted accordingly. Rather than clustering a single dataset to produce a static map of functional modules, the focus today is on data integration, network alignment, interactive visualization and ontological markup. Because of the intrinsic noisiness of high-throughput measurements, statistical methods have been central to this effort. In this review, we briefly survey available datasets in functional genomics, review methods for data integration and network alignment, and describe recent work on using network models to guide experimental validation. We explain how the integration and validation steps spring from a Bayesian description of network uncertainty, and conclude by describing an important near-term milestone for systems biology: the construction of a set of rich reference networks for key model organisms.
Balaji S. Srinivasan, Nigam H. Shah, Jason Flannick, Eduardo Abeliuk, Antal F. Novak, Serafim Batzoglou
Briefings Bioinform.2
2007 Annotation and query of tissue microarray data using the NCI Thesaurus
abstract
BACKGROUND: The Stanford Tissue Microarray Database (TMAD) is a repository of data serving a consortium of pathologists and biomedical researchers. The tissue samples in TMAD are annotated with multiple free-text fields, specifying the pathological diagnoses for each sample. These text annotations are not structured according to any ontology, making future integration of this resource with other biological and clinical data difficult. RESULTS: We developed methods to map these annotations to the NCI thesaurus. Using the NCI-T we can effectively represent annotations for about 86% of the samples. We demonstrate how this mapping enables ontology driven integration and querying of tissue microarray data. We have deployed the mapping and ontology driven querying tools at the TMAD site for general use. CONCLUSION: We have demonstrated that we can effectively map the diagnosis-related terms describing a sample in TMAD to the NCI-T. The NCI thesaurus terms have a wide coverage and provide terms for about 86% of the samples. In our opinion the NCI thesaurus can facilitate integration of this resource with other biological data.
Nigam H. Shah, Daniel L. Rubin, Inigo Espinosa, Kelli Montgomery, Mark A. Musen
BMC Bioinform.1
2006 Ontology-based Annotation and Query of Tissue Microarray Data
Nigam H. Shah, Daniel L. Rubin, Kaustubh Supekar, Mark A. Musen
AMIA1
2006 A case study in pathway knowledgebase verification
abstract
BACKGROUND: Biological databases and pathway knowledge-bases are proliferating rapidly. We are developing software tools for computer-aided hypothesis design and evaluation, and we would like our tools to take advantage of the information stored in these repositories. But before we can reliably use a pathway knowledge-base as a data source, we need to proofread it to ensure that it can fully support computer-aided information integration and inference. RESULTS: We design a series of logical tests to detect potential problems we might encounter using a particular knowledge-base, the Reactome database, with a particular computer-aided hypothesis evaluation tool, HyBrow. We develop an explicit formal language from the language implicit in the Reactome data format and specify a logic to evaluate models expressed using this language. We use the formalism of finite model theory in this work. We then use this logic to formulate tests for desirable properties (such as completeness, consistency, and well-formedness) for pathways stored in Reactome. We apply these tests to the publicly available Reactome releases (releases 10 through 14) and compare the results, which highlight Reactome's steady improvement in terms of decreasing inconsistencies. We also investigate and discuss Reactome's potential for supporting computer-aided inference tools. CONCLUSION: The case study described in this work demonstrates that it is possible to use our model theory based approach to identify problems one might encounter using a knowledge-base to support hypothesis evaluation tools. The methodology we use is general and is in no way restricted to the specific knowledge-base employed in this case study. Future application of this methodology will enable us to compare pathway resources with respect to the generic properties such resources will need to possess if they are to support automated reasoning.
Stephen A. Racunas, Nigam H. Shah, Nina V. Fedoroff
BMC Bioinform.2
2004 CLENCH: a program for calculating Cluster ENriCHment using the Gene Ontology
abstract
SUMMARY: Analysis of microarray data most often produces lists of genes with similar expression patterns, which are then subdivided into functional categories for biological interpretation. Such functional categorization is most commonly accomplished using Gene Ontology (GO) categories. Although there are several programs that identify and analyze functional categories for human, mouse and yeast genes, none of them accept Arabidopsis thaliana data. In order to address this need for A.thaliana community, we have developed a program that retrieves GO annotations for A.thaliana genes and performs functional category analysis for lists of genes selected by the user. AVAILABILITY: http://www.personal.psu.edu/nhs109/Clench
Nigam H. Shah, Nina V. Fedoroff
Bioinform.1
2003 A tool-kit for cDNA microarray and promoter analysis
abstract
We describe two sets of programs for expediting routine tasks in analysis of cDNA microarray data and promoter sequences. The first set permits bad data points to be flagged with respect to a number of parameters and performs normalization in three different ways. It allows combining of result files into comprehensive data sets, evaluation of the quality of both technical and biological replicates and row and/or column standardization of data matrices. The second set supports mapping ESTs in the genome, identifying the corresponding genes and recovering their promoters, analyzing promoters for transcription factor binding sites, and visual representation of the results. The programs are designed primarily for Arabidopsis thaliana researchers, but can be adapted readily for other model systems. Availability and Supplementary information: http://www.personal.psu.edu/nhs109/Programs/
Nigam H. Shah, D. C. King, P. N. Shah, Nina V. Fedoroff
Bioinform.1