Jeremy L. Warner

dblp:131/4621 · DBLP profile ↗
← Back
40ranked-venue papers
13as first author
6since 2021 · last 2025
0000-0002-2851-7242ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 13 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Using Large Language Model for Efficient Extraction of Treatment Discontinuation Information - A Study of Online Breast Cancer Community Posts
Qingyuan Song, Jessie Yang, Ndidiamaka Obi, Congning Ni, Jeremy L. Warner, Qingxia Chen, S. Trent Rosenbloom, Bradley A. Malin, Zhijun Yin
AIME (2)5
2025 Collaborative large language models for automated data extraction in living systematic reviews
abstract
OBJECTIVE: Data extraction from the published literature is the most laborious step in conducting living systematic reviews (LSRs). We aim to build a generalizable, automated data extraction workflow leveraging large language models (LLMs) that mimics the real-world 2-reviewer process. MATERIALS AND METHODS: A dataset of 10 trials (22 publications) from a published LSR was used, focusing on 23 variables related to trial, population, and outcomes data. The dataset was split into prompt development (n = 5) and held-out test sets (n = 17). GPT-4-turbo and Claude-3-Opus were used for data extraction. Responses from the 2 LLMs were considered concordant if they were the same for a given variable. The discordant responses from each LLM were provided to the other LLM for cross-critique. Accuracy, ie, the total number of correct responses divided by the total number of responses, was computed to assess performance. RESULTS: In the prompt development set, 110 (96%) responses were concordant, achieving an accuracy of 0.99 against the gold standard. In the test set, 342 (87%) responses were concordant. The accuracy of the concordant responses was 0.94. The accuracy of the discordant responses was 0.41 for GPT-4-turbo and 0.50 for Claude-3-Opus. Of the 49 discordant responses, 25 (51%) became concordant after cross-critique, increasing accuracy to 0.76. DISCUSSION: Concordant responses by the LLMs are likely to be accurate. In instances of discordant responses, cross-critique can further increase the accuracy. CONCLUSION: Large language models, when simulated in a collaborative, 2-reviewer workflow, can extract data with reasonable performance, enabling truly "living" systematic reviews.
Umair Ayub, Syed Arsalan Ahmed Naqvi, Kaneez Zahra Rubab Khakwani, Zaryab bin Riaz Sipra, Ammad Raina, Sihan Zhou, Amir Saeidi, Bashar Hasan, Robert Bryan Rumble, Danielle S. Bitterman, Jeremy L. Warner, Jia Zou 0001, Amye J. Tevaarwerk, Konstantinos Leventakos, Kenneth L. Kehl, Jeanne M. Palmer, Mohammad Hassan Murad, Chitta Baral, Irbaz Bin Riaz
J. Am. Medical Informatics Assoc.13
2023 Next-generation phenotyping: introducing phecodeX for enhanced discovery research in medical phenomics
abstract
MOTIVATION: Phecodes are widely used and easily adapted phenotypes based on International Classification of Diseases codes. The current version of phecodes (v1.2) was designed primarily to study common/complex diseases diagnosed in adults; however, there are numerous limitations in the codes and their structure. RESULTS: Here, we present phecodeX, an expanded version of phecodes with a revised structure and 1,761 new codes. PhecodeX adds granularity to phenotypes in key disease domains that are under-represented in the current phecode structure-including infectious disease, pregnancy, congenital anomalies, and neonatology-and is a more robust representation of the medical phenome for global use in discovery research. AVAILABILITY AND IMPLEMENTATION: phecodeX is available at https://github.com/PheWAS/phecodeX.
Megan M. Shuey, William W. Stead, Ida Aka, April L. Barnado, Lisa Bastarache, Elly Brokamp, Meredith Campbell, Robert J. Carroll, Jeffrey A. Goldstein, Adam Lewis, Beth A. Malow, Jonathan D. Mosley, Travis Osterman, Dolly A Padovani-Claudio, Andrea Ramirez, Dan M. Roden, Bryce A. Schuler, Edward Siew, Jennifer Sucre, Isaac Thomsen, Rory J. Tinker, Sara Van Driest, Colin Walsh, Jeremy L. Warner, Quinn Stanton Wells, Lee E. Wheless
Bioinform.24
2022 DeepPhe: Natural Language Processing Tools for Cancer Research and Surveillance
Harry Hochheiser, Sean Finan, Zhou Yuan, John D. Levander, Eric B. Durbin, Isaac Hands, Ramakanth Kavuluru, Jeremy L. Warner, Guergana K. Savova
AMIA8
2022 Quantitating and assessing interoperability between electronic health records
abstract
OBJECTIVES: Electronic health records (EHRs) contain a large quantity of machine-readable data. However, institutions choose different EHR vendors, and the same product may be implemented differently at different sites. Our goal was to quantify the interoperability of real-world EHR implementations with respect to clinically relevant structured data. MATERIALS AND METHODS: We analyzed de-identified and aggregated data from 68 oncology sites that implemented 1 of 5 EHR vendor products. Using 6 medications and 6 laboratory tests for which well-accepted standards exist, we calculated inter- and intra-EHR vendor interoperability scores. RESULTS: The mean intra-EHR vendor interoperability score was 0.68 as compared to a mean of 0.22 for inter-system interoperability, when weighted by number of systems of each type, and 0.57 and 0.20 when not weighting by number of systems of each type. DISCUSSION: In contrast to data elements required for successful billing, clinically relevant data elements are rarely standardized, even though applicable standards exist. We chose a representative sample of laboratory tests and medications for oncology practices, but our set of data elements should be seen as an example, rather than a definitive list. CONCLUSIONS: We defined and demonstrated a quantitative measure of interoperability between site EHR systems and within/between implemented vendor systems. Two sites that share the same vendor are, on average, more interoperable. However, even for implementation of the same EHR product, interoperability is not guaranteed. Our results can inform institutional EHR selection, analysis, and optimization for interoperability.
Elmer V. Bernstam, Jeremy L. Warner, John C. Krauss, Edward P. Ambinder, Wendy S. Rubinstein, George Komatsoulis, Robert S. Miller, James L. Chen
J. Am. Medical Informatics Assoc.2
2021 A retrospective approach to evaluating potential adverse outcomes associated with delay of procedures for cardiovascular and cancer-related diagnoses in the context of COVID-19
Neil S. Zheng, Jeremy L. Warner, Travis Osterman, Quinn Stanton Wells, Xiao-Ou Shu, Steve Deppen, Seth J. Karp, Shon Dwyer, QiPing Feng, Nancy J. Cox, Josh F. Peterson, C. Michael Stein, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics2
2020 Recommendations for patient similarity classes: results of the AMIA 2019 workshop on defining patient similarity
abstract
Defining patient-to-patient similarity is essential for the development of precision medicine in clinical care and research. Conceptually, the identification of similar patient cohorts appears straightforward; however, universally accepted definitions remain elusive. Simultaneously, an explosion of vendors and published algorithms have emerged and all provide varied levels of functionality in identifying patient similarity categories. To provide clarity and a common framework for patient similarity, a workshop at the American Medical Informatics Association 2019 Annual Meeting was convened. This workshop included invited discussants from academics, the biotechnology industry, the FDA, and private practice oncology groups. Drawing from a broad range of backgrounds, workshop participants were able to coalesce around 4 major patient similarity classes: (1) feature, (2) outcome, (3) exposure, and (4) mixed-class. This perspective expands into these 4 subtypes more critically and offers the medical informatics community a means of communicating their work on this important topic.
Nathan D. Seligson, Jeremy L. Warner, William S. Dalton, Robert S. Miller, Debra Patt, Kenneth L. Kehl, Matvey Palchuk, Gil Alterovitz, Laura K. Wiley, Ming Huang 0006, Feichen Shen, Yanshan Wang, Khoa A. Nguyen, Anthony F. Wong, Funda Meric-Bernstam, Elmer V. Bernstam, James L. Chen
J. Am. Medical Informatics Assoc.2
2019 Continuum of Interoperability in Oncology EHR Implementations
Elmer V. Bernstam, Jeremy L. Warner, John C. Krauss, Edward P. Ambinder, Wendy S. Rubinstein, George Komatsoulis, Robert S. Miller, James L. Chen
AMIA2
2019 Flipping Clinical Documentation on its Head
Yaa A. Kumah-Crystal, Liz Salmi, Chethan Sarabu, Jeffery Smith, Jeremy L. Warner
AMIA5
2019 HemOnc.org: Evaluation of Information Models for Cancer Therapy Representation
Zachary H. Moldwin, Harry Hochheiser, Jeremy L. Warner
AMIA3
2019 Patient Messaging Content Associated with Initiating Hormonal Therapy after a Breast Cancer Diagnosis
Zhijun Yin, Jeremy L. Warner, Qingxia Chen, Bradley A. Malin
AMIA2
2019 HemOnc: A new standard vocabulary for chemotherapy regimen representation in the OMOP common data model
abstract
Systematic application of observational data to the understanding of impacts of cancer treatments requires detailed information models allowing meaningful comparisons between treatment regimens. Unfortunately, details of systemic therapies are scarce in registries and data warehouses, primarily due to the complex nature of the protocols and a lack of standardization. Since 2011, we have been creating a curated and semi-structured website of chemotherapy regimens, HemOnc.org. In coordination with the Observational Health Data Sciences and Informatics (OHDSI) Oncology Subgroup, we have transformed a substantial subset of this content into the OMOP common data model, with bindings to multiple external vocabularies, e.g., RxNorm and the National Cancer Institute Thesaurus. Currently, there are >73,000 concepts and >177,000 relationships in the full vocabulary. Content related to the definition and composition of chemotherapy regimens has been released within the ATHENA tool (athena.ohdsi.org) for widespread utilization by the OHDSI membership. Here, we describe the rationale, data model, and initial contents of the HemOnc vocabulary along with several use cases for which it may be valuable.
Jeremy L. Warner, Dmitry Dymshyts, Christian G. Reich, Michael J. Gurley, Harry Hochheiser, Zachary H. Moldwin, Rimma Belenkaya, Andrew E. Williams, Peter C. Yang
J. Biomed. Informatics1
2018 connecTPL: A Tool for Connecting Drugs and Publications to their Clinical Trials
Krysten Harvey, Andrew M. Malty, Samuel M. Rubinstein, Jeremy L. Warner
AMIA4
2018 Exploring Mutational Heterogeneity in the GENIE Tumor Database Using Visual Analytics
Jeremy L. Warner, Samuel M. Rubinstein, Julie Wu, Emmanuel A. Santillana, Suresh K. Bhavnani
AMIA1
2018 Computable Longitudinal Patient Trajectories
Jeremy L. Warner, Guergana K. Savova, Noémie Elhadad, Lisa Bastarache, David Gotz
AMIA1
2018 Learning When Communications Between Healthcare Providers Indicate Hormonal Therapy Medication Discontinuation
Zhijun Yin, Jeremy L. Warner, Bradley A. Malin
AMIA2
2018 The therapy is making me sick: how online portal communications between breast cancer patients and physicians indicate medication discontinuation
abstract
Objective: Online platforms have created a variety of opportunities for breast patients to discuss their hormonal therapy, a long-term adjuvant treatment to reduce the chance of breast cancer occurrence and mortality. The goal of this investigation is to ascertain the extent to which the messages breast cancer patients communicated through an online portal can indicate their potential for discontinuing hormonal therapy. Materials and Methods: We studied the de-identified electronic medical records of 1106 breast cancer patients who were prescribed hormonal therapy at Vanderbilt University Medical Center over a 12-year period. We designed a data-driven approach to investigate patients' patterns of messaging with healthcare providers, the topics they communicated, and the extent to which these messaging behaviors associate with the likelihood that a patient will discontinue a prescribed 5-year regimen of therapy. Results: The results indicates that messaging rate over time [hazard ratio (HR) = 1.373, P = 0.002], mentions of side effects (HR = 1.214, P = 0.006), and surgery-related topics (HR = 1.170, P = 0.034) were associated with increased risk of early medication discontinuation. In contrast, seeking professional suggestions (HR = 0.766, P = 0.002), expressing gratitude to healthcare providers (HR = 0.872, P = 0.044), and mentions of drugs used to treat side effects (HR = 0.807, P = 0.013) were associated with decreased risk of medication discontinuation. Discussion and Conclusion: This investigation suggests that patient-generated content can inform the study of health-related behaviors. Given that approximately 50% of breast cancer patients do not complete a course of hormonal therapy as described, the identification of factors associated with medication discontinuation can facilitate real-time interventions to prevent early discontinuation.
Zhijun Yin, Morgan Harrell, Jeremy L. Warner, Qingxia Chen, Daniel Fabbri, Bradley A. Malin
J. Am. Medical Informatics Assoc.3
2017 Identification of Patient Subgroups in Metastatic Breast Cancer Patients Based on Somatic Copy Number Alterations: A Bipartite Network Analysis
Suresh K. Bhavnani, Archana Ayyaswami, Tianlong Chen 0004, Jeremy L. Warner
AMIA4
2017 The Power of the Patient Voice: Learning Indicators of Treatment Adherence From An Online Breast Cancer Forum
Zhijun Yin, Bradley A. Malin, Jeremy L. Warner, Pei-Yun Sabrina Hsueh, Ching-Hua Chen
ICWSM3
2016 Workshop on Visual Analytics in Healthcare
Jesus J. Caban, Adam Perer, Uba Backonja, Jeremy L. Warner, Shira Fischer
AMIA4
2016 Pragmatic precision oncology: the secondary uses of clinical tumor molecular profiling
abstract
BACKGROUND: Precision oncology increasingly utilizes molecular profiling of tumors to determine treatment decisions with targeted therapeutics. The molecular profiling data is valuable in the treatment of individual patients as well as for multiple secondary uses. OBJECTIVE: To automatically parse, categorize, and aggregate clinical molecular profile data generated during cancer care as well as use this data to address multiple secondary use cases. METHODS: A system to parse, categorize and aggregate molecular profile data was created. A naÿve Bayesian classifier categorized results according to clinical groups. The accuracy of these systems were validated against a published expertly-curated subset of molecular profiling data. RESULTS: Following one year of operation, 819 samples have been accurately parsed and categorized to generate a data repository of 10,620 genetic variants. The database has been used for operational, clinical trial, and discovery science research. CONCLUSIONS: A real-time database of molecular profiling data is a pragmatic solution to several knowledge management problems in the practice and science of precision oncology.
Matthew J. Rioth, Ramya Thota, David B. Staggs, Douglas B. Johnson, Jeremy L. Warner
J. Am. Medical Informatics Assoc.5
2016 SMART precision cancer medicine: a FHIR-based app to provide genomic information at the point of care
abstract
BACKGROUND: Precision cancer medicine (PCM) will require ready access to genomic data within the clinical workflow and tools to assist clinical interpretation and enable decisions. Since most electronic health record (EHR) systems do not yet provide such functionality, we developed an EHR-agnostic, clinico-genomic mobile app to demonstrate several features that will be needed for point-of-care conversations. METHODS: Our prototype, called Substitutable Medical Applications and Reusable Technology (SMART)® PCM, visualizes genomic information in real time, comparing a patient's diagnosis-specific somatic gene mutations detected by PCR-based hotspot testing to a population-level set of comparable data. The initial prototype works for patient specimens with 0 or 1 detected mutation. Genomics extensions were created for the Health Level Seven® Fast Healthcare Interoperability Resources (FHIR)® standard; otherwise, the prototype is a normal SMART on FHIR app. RESULTS: The PCM prototype can rapidly present a visualization that compares a patient's somatic genomic alterations against a distribution built from more than 3000 patients, along with context-specific links to external knowledge bases. Initial evaluation by oncologists provided important feedback about the prototype's strengths and weaknesses. We added several requested enhancements and successfully demonstrated the app at the inaugural American Society of Clinical Oncology Interoperability Demonstration; we have also begun to expand visualization capabilities to include cancer specimens with multiple mutations. DISCUSSION: PCM is open-source software for clinicians to present the individual patient within the population-level spectrum of cancer somatic mutations. The app can be implemented on any SMART on FHIR-enabled EHRs, and future versions of PCM should be able to evolve in parallel with external knowledge bases.
Jeremy L. Warner, Matthew J. Rioth, Kenneth D. Mandl, Joshua C. Mandel, David A. Kreda, Isaac S. Kohane, Daniel Carbone, Ross Oreto, Lucy Wang, Shilin Zhu, Heming Yao, Gil Alterovitz
J. Am. Medical Informatics Assoc.1
2016 CUSTOM-SEQ: a prototype for oncology rapid learning in a comprehensive EHR environment
abstract
BACKGROUND: As targeted cancer therapies and molecular profiling become widespread, the era of "precision oncology" is at hand. However, cancer genomes are complex, making mutation-specific outcomes difficult to track. We created a proof-of-principle, CUSTOM-SEQ: Continuously Updating System for Tracking Outcome by Mutation, to Support Evidence-based Querying, to automatically calculate and display mutation-specific survival statistics from electronic health record data. METHODS: Patients with cancer genotyping were included, and clinical data was extracted through a variety of algorithms. Results were refreshed regularly and injected into a standard reporting platform. Significant results were highlighted for visual cueing. A subset was additionally stratified by stage, smoking status, and treatment exposure. RESULTS: By August 2015, 4310 patients with a median follow-up of 17 months had sufficient data for survival calculation. As expected, epidermal growth factor receptor (EGFR) mutations in lung cancer were associated with superior overall survival, hazard ratio (HR) = 0.53 (P < .001), validating the approach. Guanine nucleotide binding protein (G protein), q polypeptide (GNAQ) mutations in melanoma were associated with inferior overall survival, a novel finding (HR = 3.42, P < .001). Smoking status was not prognostic for epidermal growth factor receptor-mutated lung cancer patients, who also lived significantly longer than their counterparts, even with advanced disease (HR = 0.54, P = .001). INTERPRETATION: CUSTOM-SEQ represents a novel rapid learning system for a precision oncology environment. Retrospective studies are often limited by study of specific time periods and can lead to incomplete conclusions. Because data is continuously updated in CUSTOM-SEQ, the evidence base is constantly growing. Future work will allow users to interactively explore populations by demographics and treatment exposure, in order to further investigate significant mutation-specific signals.
Jeremy L. Warner, Lucy Wang, William Pao, Jeffrey A. Sosman, Ravi V. Atreya, Pam Carney, Mia A. Levy
J. Am. Medical Informatics Assoc.1
2016 Combining billing codes, clinical notes, and medications from electronic health records provides superior phenotyping performance
abstract
OBJECTIVE: To evaluate the phenotyping performance of three major electronic health record (EHR) components: International Classification of Disease (ICD) diagnosis codes, primary notes, and specific medications. MATERIALS AND METHODS: We conducted the evaluation using de-identified Vanderbilt EHR data. We preselected ten diseases: atrial fibrillation, Alzheimer's disease, breast cancer, gout, human immunodeficiency virus infection, multiple sclerosis, Parkinson's disease, rheumatoid arthritis, and types 1 and 2 diabetes mellitus. For each disease, patients were classified into seven categories based on the presence of evidence in diagnosis codes, primary notes, and specific medications. Twenty-five patients per disease category (a total number of 175 patients for each disease, 1750 patients for all ten diseases) were randomly selected for manual chart review. Review results were used to estimate the positive predictive value (PPV), sensitivity, andF-score for each EHR component alone and in combination. RESULTS: The PPVs of single components were inconsistent and inadequate for accurately phenotyping (0.06-0.71). Using two or more ICD codes improved the average PPV to 0.84. We observed a more stable and higher accuracy when using at least two components (mean ± standard deviation: 0.91 ± 0.08). Primary notes offered the best sensitivity (0.77). The sensitivity of ICD codes was 0.67. Again, two or more components provided a reasonably high and stable sensitivity (0.59 ± 0.16). Overall, the best performance (Fscore: 0.70 ± 0.12) was achieved by using two or more components. Although the overall performance of using ICD codes (0.67 ± 0.14) was only slightly lower than using two or more components, its PPV (0.71 ± 0.13) is substantially worse (0.91 ± 0.08). CONCLUSION: Multiple EHR components provide a more consistent and higher performance than a single one for the selected phenotypes. We suggest considering multiple EHR components for future phenotyping design in order to obtain an ideal result.
Wei-Qi Wei, Pedro L. Teixeira, Huan Mo, Robert M. Cronin, Jeremy L. Warner, Joshua C. Denny
J. Am. Medical Informatics Assoc.5
2016 Classification of hospital acquired complications using temporal clinical information from a large electronic health record
Jeremy L. Warner, Peijin Zhang, Jenny Liu, Gil Alterovitz
J. Biomed. Informatics1
2015 Visualizing High Dimensional Clinical and Tumor Genotyping Data
Matthew J. Rioth, Jeremy L. Warner
AMIA2
2015 SMART on FHIR Genomics: facilitating standardized clinico-genomic apps
abstract
BACKGROUND: Supporting clinical decision support for personalized medicine will require linking genome and phenome variants to a patient's electronic health record (EHR), at times on a vast scale. Clinico-genomic data standards will be needed to unify how genomic variant data are accessed from different sequencing systems. METHODS: A specification for the basis of a clinic-genomic standard, building upon the current Health Level Seven International Fast Healthcare Interoperability Resources (FHIR®) standard, was developed. An FHIR application protocol interface (API) layer was attached to proprietary sequencing platforms and EHRs in order to expose gene variant data for presentation to the end-user. Three representative apps based on the SMART platform were built to test end-to-end feasibility, including integration of genomic and clinical data. RESULTS: Successful design, deployment, and use of the API was demonstrated and adopted by HL7 Clinical Genomics Workgroup. Feasibility was shown through development of three apps by various types of users with background levels and locations. CONCLUSION: This prototyping work suggests that an entirely data (and web) standards-based approach could prove both effective and efficient for advancing personalized medicine.
Gil Alterovitz, Jeremy L. Warner, Peijin Zhang, Yishen Chen, Mollie Ullman-Cullere, David A. Kreda, Isaac S. Kohane
J. Am. Medical Informatics Assoc.2
2015 Seeing the forest through the trees: uncovering phenomic complexity through interactive network visualization
abstract
Our aim was to uncover unrecognized phenomic relationships using force-based network visualization methods, based on observed electronic medical record data. A primary phenotype was defined from actual patient profiles in the Multiparameter Intelligent Monitoring in Intensive Care II database. Network visualizations depicting primary relationships were compared to those incorporating secondary adjacencies. Interactivity was enabled through a phenotype visualization software concept: the Phenomics Advisor. Subendocardial infarction with cardiac arrest was demonstrated as a sample phenotype; there were 332 primarily adjacent diagnoses, with 5423 relationships. Primary network visualization suggested a treatment-related complication phenotype and several rare diagnoses; re-clustering by secondary relationships revealed an emergent cluster of smokers with the metabolic syndrome. Network visualization reveals phenotypic patterns that may have remained occult in pairwise correlation analysis. Visualization of complex data, potentially offered as point-of-care tools on mobile devices, may allow clinicians and researchers to quickly generate hypotheses and gain deeper understanding of patient subpopulations.
Jeremy L. Warner, Joshua C. Denny, David A. Kreda, Gil Alterovitz
J. Am. Medical Informatics Assoc.1
2015 Development, implementation, and initial evaluation of a foundational open interoperability standard for oncology treatment planning and summarization
abstract
OBJECTIVE: Develop and evaluate a foundational oncology-specific standard for the communication and coordination of care throughout the cancer journey, with early-stage breast cancer as the use case. MATERIALS AND METHODS: Owing to broad uptake of the Health Level Seven (HL7) Consolidated Clinical Document Architecture (C-CDA) by health information exchanges and large provider organizations, we developed an implementation guide in congruence with C-CDA. The resultant product was balloted through the HL7 process and subsequently implemented by two groups: the Health Story Project (Health Story) and the Athena Breast Health Network (Athena). RESULTS: The HL7 Implementation Guide for CDA, Release 2: Clinical Oncology Treatment Plan and Summary, DSTU Release 1 (eCOTPS) was successfully balloted and published as a Draft Standard for Trial Use (DSTU) in October 2013. Health Story successfully implemented the eCOTPS the 2014 meeting of the Healthcare Information and Management Systems Society (HIMSS) in a clinical vignette. During the evaluation and implementation of eCOPS, Athena identified two practical concerns: (1) the need for additional CDA templates specific to their use case; (2) the many-to-many mapping of Athena-defined data elements to eCOTPS. DISCUSSION: Early implementation of eCOTPS has demonstrated successful vendor-agnostic transmission of oncology-specific data. The modularity enabled by the C-CDA framework ensures the relatively straightforward expansion of the eCOTPS to include other cancer subtypes. Lessons learned during the process will strengthen future versions of the standard. CONCLUSION: eCOTPS is the first oncology-specific CDA standard to achieve HL7 DSTU status. Oncology standards will improve care throughout the cancer journey by allowing the efficient transmission of reliable, meaningful, and current clinical data between the many involved stakeholders.
Jeremy L. Warner, Suzanne E. Maddux, Kevin S. Hughes, John C. Krauss, Peter Paul Yu, Lawrence N. Shulman, Deborah K. Mayer, Michael A. Hogarth, Mark Shafarman, Allison Stover Fiscalini, Laura Esserman, Liora Alschuler, George A. Koromia, Zabrina Gonzaga, Edward P. Ambinder
J. Am. Medical Informatics Assoc.1
2015 Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortality
abstract
OBJECTIVES: Drug repurposing, which finds new indications for existing drugs, has received great attention recently. The goal of our work is to assess the feasibility of using electronic health records (EHRs) and automated informatics methods to efficiently validate a recent drug repurposing association of metformin with reduced cancer mortality. METHODS: By linking two large EHRs from Vanderbilt University Medical Center and Mayo Clinic to their tumor registries, we constructed a cohort including 32,415 adults with a cancer diagnosis at Vanderbilt and 79,258 cancer patients at Mayo from 1995 to 2010. Using automated informatics methods, we further identified type 2 diabetes patients within the cancer cohort and determined their drug exposure information, as well as other covariates such as smoking status. We then estimated HRs for all-cause mortality and their associated 95% CIs using stratified Cox proportional hazard models. HRs were estimated according to metformin exposure, adjusted for age at diagnosis, sex, race, body mass index, tobacco use, insulin use, cancer type, and non-cancer Charlson comorbidity index. RESULTS: Among all Vanderbilt cancer patients, metformin was associated with a 22% decrease in overall mortality compared to other oral hypoglycemic medications (HR 0.78; 95% CI 0.69 to 0.88) and with a 39% decrease compared to type 2 diabetes patients on insulin only (HR 0.61; 95% CI 0.50 to 0.73). Diabetic patients on metformin also had a 23% improved survival compared with non-diabetic patients (HR 0.77; 95% CI 0.71 to 0.85). These associations were replicated using the Mayo Clinic EHR data. Many site-specific cancers including breast, colorectal, lung, and prostate demonstrated reduced mortality with metformin use in at least one EHR. CONCLUSIONS: EHR data suggested that the use of metformin was associated with decreased mortality after a cancer diagnosis compared with diabetic and non-diabetic cancer patients not on metformin, indicating its potential as a chemotherapeutic regimen. This study serves as a model for robust and inexpensive validation studies for drug repurposing signals using EHR data.
Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Xiaoyang Ruan, Min Jiang 0007, Jamii St Julien, Jeremy L. Warner, Carol Friedman, Dan M. Roden, Joshua C. Denny
J. Am. Medical Informatics Assoc.14
2014 PheWAS and Genetics Define Subphenotypes in Drug Response
Robert J. Carroll, Jeremy L. Warner, Anne E. Eyler, Charles Moore, Jayanth Doss, Katherine P. Liao, Robert M. Plenge, Joshua C. Denny
AMIA2
2014 Identifying Metastases from Pathology Reports in Lung Cancer Patients
Ergin Soysal, Jeremy L. Warner, Joshua C. Denny, Hua Xu 0001
AMIA2
2014 Evaluation of Diagnosis Codes, Clinical Notes, and Medications on Identifying Subjects with a Specific Disease Phenotype
Wei-Qi Wei, Pedro L. Teixeira, Huan Mo, Robert M. Cronin, Jeremy L. Warner, Joshua C. Denny
AMIA5
2014 Mining electronic health record data to detect drug-repurposing signals for cancers
Hua Xu 0001, Qingxia Chen, Jeremy L. Warner, Min Jiang 0007, Anushi Shah, Melinda Aldrich, Joshua C. Denny
AMIA3
2013 Constructing a Novel Cancer Ontology
Michael C. Gao, Jeremy L. Warner, Peter C. Yang, Gil Alterovitz
AMIA2
2013 Sharing of Genomic Information: Perspectives from Stakeholders
Jeremy L. Warner, Gil Alterovitz, Joshua C. Denny, Robert Fassett, Kevin S. Hughes
AMIA1
2013 Phenometric analysis of electronic health records: a new approach to visualization of high dimensional biomedical information
Jeremy L. Warner, Quan Ding, David A. Kreda, Zi'ou Zheng, Joshua C. Denny, Gil Alterovitz
AMIA1
2013 Brief communication: External phenome analysis enables a rational federated query strategy to detect changing rates of treatment-related complications associated with multiple myeloma
abstract
Electronic health records (EHRs) are increasingly useful for health services research. For relatively uncommon conditions, such as multiple myeloma (MM) and its treatment-related complications, a combination of multiple EHR sources is essential for such research. The Shared Health Research Information Network (SHRINE) enables queries for aggregate results across participating institutions. Development of a rational search strategy in SHRINE may be augmented through analysis of pre-existing databases. We developed a SHRINE query for likely non-infectious treatment-related complications of MM, based upon an analysis of the Multiparameter Intelligent Monitoring in Intensive Care (MIMIC II) database. Using this query strategy, we found that the rate of likely treatment-related complications significantly increased from 2001 to 2007, by an average of 6% a year (p=0.01), across the participating SHRINE institutions. This finding is in keeping with increasingly aggressive strategies in the treatment of MM. This proof of concept demonstrates that a staged approach to federated queries, using external EHR data, can yield potentially clinically meaningful results.
Jeremy L. Warner, Gil Alterovitz, Kelly Bodio, Robin M. Joyce
J. Am. Medical Informatics Assoc.1
2012 Phenome-Based Analysis as a Means for Discovering Context-Dependent Clinical Reference Ranges
Jeremy L. Warner, Gil Alterovitz
AMIA1
2012 Towards an Annotation Schema for Cancer Trajectory State Detection
Jeremy L. Warner, Peter Anick, Kenneth Roach, Nianwen Xue, Robin M. Joyce, Charles Safran, Pengyu Hong
AMIA1