Josh F. Peterson

dblp:07/7342 · DBLP profile ↗
← Back
40ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0002-7553-0749ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 40 · 5 first-author · 12 since 2021
YearPublicationVenuePosition
2024 Leveraging explainable artificial intelligence to optimize clinical decision support
abstract
OBJECTIVE: To develop and evaluate a data-driven process to generate suggestions for improving alert criteria using explainable artificial intelligence (XAI) approaches. METHODS: We extracted data on alerts generated from January 1, 2019 to December 31, 2020, at Vanderbilt University Medical Center. We developed machine learning models to predict user responses to alerts. We applied XAI techniques to generate global explanations and local explanations. We evaluated the generated suggestions by comparing with alert's historical change logs and stakeholder interviews. Suggestions that either matched (or partially matched) changes already made to the alert or were considered clinically correct were classified as helpful. RESULTS: The final dataset included 2 991 823 firings with 2689 features. Among the 5 machine learning models, the LightGBM model achieved the highest Area under the ROC Curve: 0.919 [0.918, 0.920]. We identified 96 helpful suggestions. A total of 278 807 firings (9.3%) could have been eliminated. Some of the suggestions also revealed workflow and education issues. CONCLUSION: We developed a data-driven process to generate suggestions for improving alert criteria using XAI techniques. Our approach could identify improvements regarding clinical decision support (CDS) that might be overlooked or delayed in manual reviews. It also unveils a secondary purpose for the XAI: to improve quality by discovering scenarios where CDS alerts are not accepted due to workflow, education, or staffing issues.
Siru Liu, Allison B. McCoy, Josh F. Peterson, Thomas A. Lasko, Dean F. Sittig, Scott D. Nelson, Jennifer Andrews, Lorraine Patterson, Cheryl M. Cobb, David Mulherin, Colleen T. Morton, Adam Wright
J. Am. Medical Informatics Assoc.3
2024 Leveraging large language models for generating responses to patient messages - a subjective analysis
abstract
OBJECTIVE: This study aimed to develop and assess the performance of fine-tuned large language models for generating responses to patient messages sent via an electronic health record patient portal. MATERIALS AND METHODS: Utilizing a dataset of messages and responses extracted from the patient portal at a large academic medical center, we developed a model (CLAIR-Short) based on a pre-trained large language model (LLaMA-65B). In addition, we used the OpenAI API to update physician responses from an open-source dataset into a format with informative paragraphs that offered patient education while emphasizing empathy and professionalism. By combining with this dataset, we further fine-tuned our model (CLAIR-Long). To evaluate fine-tuned models, we used 10 representative patient portal questions in primary care to generate responses. We asked primary care physicians to review generated responses from our models and ChatGPT and rated them for empathy, responsiveness, accuracy, and usefulness. RESULTS: The dataset consisted of 499 794 pairs of patient messages and corresponding responses from the patient portal, with 5000 patient messages and ChatGPT-updated responses from an online platform. Four primary care physicians participated in the survey. CLAIR-Short exhibited the ability to generate concise responses similar to provider's responses. CLAIR-Long responses provided increased patient educational content compared to CLAIR-Short and were rated similarly to ChatGPT's responses, receiving positive evaluations for responsiveness, empathy, and accuracy, while receiving a neutral rating for usefulness. CONCLUSION: This subjective analysis suggests that leveraging large language models to generate responses to patient messages demonstrates significant potential in facilitating communication between patients and healthcare providers.
Siru Liu, Allison B. McCoy, Aileen P. Wright, Babatunde Carew, Julian Z. Genkins, Sean S. Huang, Josh F. Peterson, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.7
2024 Using large language model to guide patients to create efficient and comprehensive clinical care message
abstract
OBJECTIVE: This study aims to investigate the feasibility of using Large Language Models (LLMs) to engage with patients at the time they are drafting a question to their healthcare providers, and generate pertinent follow-up questions that the patient can answer before sending their message, with the goal of ensuring that their healthcare provider receives all the information they need to safely and accurately answer the patient's question, eliminating back-and-forth messaging, and the associated delays and frustrations. METHODS: We collected a dataset of patient messages sent between January 1, 2022 to March 7, 2023 at Vanderbilt University Medical Center. Two internal medicine physicians identified 7 common scenarios. We used 3 LLMs to generate follow-up questions: (1) Comprehensive LLM Artificial Intelligence Responder (CLAIR): a locally fine-tuned LLM, (2) GPT4 with a simple prompt, and (3) GPT4 with a complex prompt. Five physicians rated them with the actual follow-ups written by healthcare providers on clarity, completeness, conciseness, and utility. RESULTS: For five scenarios, our CLAIR model had the best performance. The GPT4 model received higher scores for utility and completeness but lower scores for clarity and conciseness. CLAIR generated follow-up questions with similar clarity and conciseness as the actual follow-ups written by healthcare providers, with higher utility than healthcare providers and GPT4, and lower completeness than GPT4, but better than healthcare providers. CONCLUSION: LLMs can generate follow-up patient messages designed to clarify a medical question that compares favorably to those generated by healthcare providers.
Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Julian Z. Genkins, Josh F. Peterson, Yaa A. Kumah-Crystal, William Martinez, Babatunde Carew, Dara Eckerle Mize, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.6
2024 Large language models facilitate the generation of electronic health record phenotyping algorithms
abstract
OBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms.
Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei
J. Am. Medical Informatics Assoc.7
2023 Scanning the medical phenome to identify new diagnoses after recovery from COVID-19 in a US cohort
abstract
OBJECTIVE: COVID-19 survivors are at risk for long-term health effects, but assessing the sequelae of COVID-19 at large scales is challenging. High-throughput methods to efficiently identify new medical problems arising after acute medical events using the electronic health record (EHR) could improve surveillance for long-term consequences of acute medical problems like COVID-19. MATERIALS AND METHODS: We augmented an existing high-throughput phenotyping method (PheWAS) to identify new diagnoses occurring after an acute temporal event in the EHR. We then used the temporal-informed phenotypes to assess development of new medical problems among COVID-19 survivors enrolled in an EHR cohort of adults tested for COVID-19 at Vanderbilt University Medical Center. RESULTS: The study cohort included 186 105 adults tested for COVID-19 from March 5, 2020 to November 1, 2021; of which 30 088 (16.2%) tested positive. Median follow-up after testing was 412 days (IQR 274-528). Our temporal-informed phenotyping was able to distinguish phenotype chapters based on chronicity of their constituent diagnoses. PheWAS with temporal-informed phenotypes identified increased risk for 43 diagnoses among COVID-19 survivors during outpatient follow-up, including multiple new respiratory, cardiovascular, neurological, and pregnancy-related conditions. Findings were robust to sensitivity analyses, and several phenotypic associations were supported by changes in outpatient vital signs or laboratory tests from the pretesting to postrecovery period. CONCLUSION: Temporal-informed PheWAS identified new diagnoses affecting multiple organ systems among COVID-19 survivors. These findings can inform future efforts to enable longitudinal health surveillance for survivors of COVID-19 and other acute medical conditions using the EHR.
Vern Eric Kerchberger, Josh F. Peterson, Wei-Qi Wei
J. Am. Medical Informatics Assoc.2
2023 Evaluating and mitigating bias in machine learning models for cardiovascular disease prediction
Fuchen Li, Patrick Wu, Henry H. Ong, Josh F. Peterson, Wei-Qi Wei, Juan Zhao 0003
J. Biomed. Informatics4
2022 High-throughput assessment of genomic outcomes: development and validation of the HI-TAG knowledgebase
Jodell E. Linder, Ni Ketut Wilmayani, Marc S. Williams, Josh F. Peterson
AMIA5
2021 Using Genomic Association Replication Rates as an EHR Quality Measure via the Phenotype-Genotype Reference Map (PGRM)
Sarah DeLozier, Josh F. Peterson, Joshua C. Denny, Lisa Bastarache
AMIA2
2021 Mapping the Read2/CTV3 controlled clinical terminologies to Phecodes in UK Biobank primary care electronic health records: implementation and evaluation
Spiros C. Denaxas, QiPing Feng, Ghazaleh Fatemifar, Lisa Bastarache, Vern Eric Kerchberger, Aroon D. Hingorani, R. Tom Lumbers, Josh F. Peterson, Wei-Qi Wei, Harry Hemingway
AMIA9
2021 Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort
Sarah DeLozier, Sarah Bland, Melissa McPheeters, Quinn Stanton Wells, Eric Farber-Eger, Cosmin Adrian Bejan, Daniel Fabbri, S. Trent Rosenbloom, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei, Josh F. Peterson, Lisa Bastarache
J. Biomed. Informatics12
2021 ConceptWAS: A high-throughput method for early identification of COVID-19 presenting symptoms and characteristics from clinical notes
Juan Zhao 0003, Monika E. Grabowska, Vern Eric Kerchberger, Joshua C. Smith, H. Nur Eken, QiPing Feng, Josh F. Peterson, S. Trent Rosenbloom, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics7
2021 A retrospective approach to evaluating potential adverse outcomes associated with delay of procedures for cardiovascular and cancer-related diagnoses in the context of COVID-19
Neil S. Zheng, Jeremy L. Warner, Travis Osterman, Quinn Stanton Wells, Xiao-Ou Shu, Steve Deppen, Seth J. Karp, Shon Dwyer, QiPing Feng, Nancy J. Cox, Josh F. Peterson, C. Michael Stein, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics11
2019 Development of a Genomic Data Flow Framework: Results of a Survey Administered to NIH-NHGRI IGNITE and eMERGE Consortia Participants
Paul Richard Dexter, Henry H. Ong, Amanda Elsey, Gillian Bell, Nephi Walton, Wendy K. Chung, Luke V. Rasmussen, J. Kevin Hicks, Aniwaa Owusu-obeng, Stuart A. Scott, Stephen B. Ellis, Josh F. Peterson
AMIA12
2019 Extracting Drug Exposure Epochs and Drug Response Outcomes from Electronic Health Records
Andrea H. Ramirez, Yaping Shi, Elliot M. Fielstein, Jonathan S. Schildcrout, Henry H. Ong, Joshua C. Denny, Josh F. Peterson
AMIA7
2019 Pharmacogenomic clinical decision support design and multi-site process outcomes analysis in the eMERGE Network
abstract
To better understand the real-world effects of pharmacogenomic (PGx) alerts, this study aimed to characterize alert design within the eMERGE Network, and to establish a method for sharing PGx alert response data for aggregate analysis. Seven eMERGE sites submitted design details and established an alert logging data dictionary. Six sites participated in a pilot study, sharing alert response data from their electronic health record systems. PGx alert design varied, with some consensus around the use of active, post-test alerts to convey Clinical Pharmacogenetics Implementation Consortium recommendations. Sites successfully shared response data, with wide variation in acceptance and follow rates. Results reflect the lack of standardization in PGx alert design. Standards and/or larger studies will be necessary to fully understand PGx impact. This study demonstrated a method for sharing PGx alert response data and established that variation in system design is a significant barrier for multi-site analyses.
Timothy M. Herr, Josh F. Peterson, Luke V. Rasmussen, Pedro J. Caraballo, Peggy L. Peissig, Justin Starren
J. Am. Medical Informatics Assoc.2
2018 Preemptive Clinical Decision Support: Delivering Precise Information Clinicians Will Actually Use to Prevent Harm
James M. Hoffman, Josh F. Peterson, Naveen Muthu, Henry M. Dunnenberger
AMIA2
2018 Development and Implementation of Genomic Data Pipelines within US institutions: A pilot multi-site survey among NHGRI's IGNITE Genomic Medicine Sites
Josh F. Peterson, Henry H. Ong, Julie A. Lynch, J. Kevin Hicks, Paul Richard Dexter
AMIA1
2018 EHR Extraction of Longitudinal Exposure to Proton Pump Inhibitors
Andrea H. Ramirez, Elliot M. Fielstein, QiPing Feng, Henry H. Ong, Jonathan S. Schildcrout, Yaping Shi, Joshua C. Denny, Josh F. Peterson
AMIA8
2018 Empowering genomic medicine by establishing critical sequencing result data flows: the eMERGE example
abstract
The eMERGE Network is establishing methods for electronic transmittal of patient genetic test results from laboratories to healthcare providers across organizational boundaries. We surveyed the capabilities and needs of different network participants, established a common transfer format, and implemented transfer mechanisms based on this format. The interfaces we created are examples of the connectivity that must be instantiated before electronic genetic and genomic clinical decision support can be effectively built at the point of care. This work serves as a case example for both standards bodies and other organizations working to build the infrastructure required to provide better electronic clinical decision support for clinicians.
Samuel J. Aronson, Lawrence J. Babb, Darren C. Ames, Richard A. Gibbs, Eric Venner, John J. Connelly, Keith Marsolo, Chunhua Weng, Marc S. Williams, Andrea L. Hartzler, Wayne H. Liang, James D. Ralston, Emily Beth Devine, Shawn N. Murphy, Christopher G. Chute, Pedro J. Caraballo, Iftikhar J. Kullo, Robert R. Freimuth, Luke V. Rasmussen, Firas H. Wehbe, Josh F. Peterson, Jamie R. Robinson, Ken Wiley, Casey Overby Taylor
J. Am. Medical Informatics Assoc.21
2016 Implementing Pharmacogenomic Clinical Decision Support: Design and Prescriber Response in the eMERGE Network
Timothy M. Herr, Josh F. Peterson, Luke V. Rasmussen, Pedro J. Caraballo
AMIA2
2016 Developing knowledge resources to support precision medicine: principles from the Clinical Pharmacogenetics Implementation Consortium (CPIC)
abstract
To move beyond a select few genes/drugs, the successful adoption of pharmacogenomics into routine clinical care requires a curated and machine-readable database of pharmacogenomic knowledge suitable for use in an electronic health record (EHR) with clinical decision support (CDS). Recognizing that EHR vendors do not yet provide a standard set of CDS functions for pharmacogenetics, the Clinical Pharmacogenetics Implementation Consortium (CPIC) Informatics Working Group is developing and systematically incorporating a set of EHR-agnostic implementation resources into all CPIC guidelines. These resources illustrate how to integrate pharmacogenomic test results in clinical information systems with CDS to facilitate the use of patient genomic data at the point of care. Based on our collective experience creating existing CPIC resources and implementing pharmacogenomics at our practice sites, we outline principles to define the key features of future knowledge bases and discuss the importance of these knowledge resources for pharmacogenomics and ultimately precision medicine.
James M. Hoffman, Henry M. Dunnenberger, J. Kevin Hicks, Kelly E. Caudle, Michelle Whirl Carrillo, Robert R. Freimuth, Marc S. Williams, Teri E. Klein, Josh F. Peterson
J. Am. Medical Informatics Assoc.9
2015 Public Implementation Resources for Genomic Medicine
Josh F. Peterson, Marc S. Williams, Casey Overby Taylor, Robert R. Freimuth, Iftikhar J. Kullo
AMIA1
2015 National Veterans Health Administration inpatient risk stratification models for hospital-acquired acute kidney injury
abstract
OBJECTIVE: Hospital-acquired acute kidney injury (HA-AKI) is a potentially preventable cause of morbidity and mortality. Identifying high-risk patients prior to the onset of kidney injury is a key step towards AKI prevention. MATERIALS AND METHODS: A national retrospective cohort of 1,620,898 patient hospitalizations from 116 Veterans Affairs hospitals was assembled from electronic health record (EHR) data collected from 2003 to 2012. HA-AKI was defined at stage 1+, stage 2+, and dialysis. EHR-based predictors were identified through logistic regression, least absolute shrinkage and selection operator (lasso) regression, and random forests, and pair-wise comparisons between each were made. Calibration and discrimination metrics were calculated using 50 bootstrap iterations. In the final models, we report odds ratios, 95% confidence intervals, and importance rankings for predictor variables to evaluate their significance. RESULTS: The area under the receiver operating characteristic curve (AUC) for the different model outcomes ranged from 0.746 to 0.758 in stage 1+, 0.714 to 0.720 in stage 2+, and 0.823 to 0.825 in dialysis. Logistic regression had the best AUC in stage 1+ and dialysis. Random forests had the best AUC in stage 2+ but the least favorable calibration plots. Multiple risk factors were significant in our models, including some nonsteroidal anti-inflammatory drugs, blood pressure medications, antibiotics, and intravenous fluids given during the first 48 h of admission. CONCLUSIONS: This study demonstrated that, although all the models tested had good discrimination, performance characteristics varied between methods, and the random forests models did not calibrate as well as the lasso or logistic regression models. In addition, novel modifiable risk factors were explored and found to be significant.
Robert M. Cronin, Jacob P. VanHouten, Edward D. Siew, Svetlana K. Eden, Stephan D. Fihn, Christopher D. Nielson, Josh F. Peterson, Clifton R. Baker, T. Alp Ikizler, Theodore Speroff, Michael E. Matheny
J. Am. Medical Informatics Assoc.7
2015 CSER and eMERGE: current and potential state of the display of genetic information in the electronic health record
abstract
OBJECTIVE: Clinicians' ability to use and interpret genetic information depends upon how those data are displayed in electronic health records (EHRs). There is a critical need to develop systems to effectively display genetic information in EHRs and augment clinical decision support (CDS). MATERIALS AND METHODS: The National Institutes of Health (NIH)-sponsored Clinical Sequencing Exploratory Research and Electronic Medical Records & Genomics EHR Working Groups conducted a multiphase, iterative process involving working group discussions and 2 surveys in order to determine how genetic and genomic information are currently displayed in EHRs, envision optimal uses for different types of genetic or genomic information, and prioritize areas for EHR improvement. RESULTS: There is substantial heterogeneity in how genetic information enters and is documented in EHR systems. Most institutions indicated that genetic information was displayed in multiple locations in their EHRs. Among surveyed institutions, genetic information enters the EHR through multiple laboratory sources and through clinician notes. For laboratory-based data, the source laboratory was the main determinant of the location of genetic information in the EHR. The highest priority recommendation was to address the need to implement CDS mechanisms and content for decision support for medically actionable genetic information. CONCLUSION: Heterogeneity of genetic information flow and importance of source laboratory, rather than clinical content, as a determinant of information representation are major barriers to using genetic information optimally in patient care. Greater effort to develop interoperable systems to receive and consistently display genetic and/or genomic information and alert clinicians to genomic-dependent improvements to clinical care is recommended.
Brian H. Shirts, Joseph S. Salama, Samuel J. Aronson, Wendy K. Chung, Stacy W. Gray, Lucia Hindorff, Gail P. Jarvik, Sharon E. Plon, Elena M. Stoffel, Peter Tarczy-Hornoch, Eliezer M. Van Allen, Karen E. Weck, Christopher G. Chute, Robert R. Freimuth, Robert Grundmeier, Andrea L. Hartzler, Rongling Li, Peggy L. Peissig, Josh F. Peterson, Luke V. Rasmussen, Justin Starren, Marc S. Williams, Casey Overby Taylor
J. Am. Medical Informatics Assoc.19
2014 A Template for Authoring and Adapting Genomic Medicine Content in the eMERGE Infobutton Project
Casey Overby Taylor, Luke V. Rasmussen, Andrea L. Hartzler, John J. Connolly, Josh F. Peterson, RoseMary Hedberg, Robert R. Freimuth, Brian H. Shirts, Joshua C. Denny, Eric B. Larson, Christopher G. Chute, Gail P. Jarvik, James D. Ralston, Alan R. Shuldiner, Iftikhar J. Kullo, Peter Tarczy-Hornoch, Marc S. Williams
AMIA5
2013 Analyzing the Impact of Pharmacogenomics on Clinical Practice: A Visual Method
Ioana Danciu, Josh F. Peterson
AMIA2
2013 Establishing the Need for Personalized Medicine: Simvastatin Exposure Among a SLCO1B1 Variant Population
Laura K. Wiley, Josh F. Peterson, Joshua C. Denny, William S. Bush
AMIA2
2012 Focus on health information technology, electronic health records and their financial impact: A framework for evaluating the appropriateness of clinical decision support alerts and responses
abstract
OBJECTIVE: Alerting systems, a type of clinical decision support, are increasingly prevalent in healthcare, yet few studies have concurrently measured the appropriateness of alerts with provider responses to alerts. Recent reports of suboptimal alert system design and implementation highlight the need for better evaluation to inform future designs. The authors present a comprehensive framework for evaluating the clinical appropriateness of synchronous, interruptive medication safety alerts. METHODS: Through literature review and iterative testing, metrics were developed that describe successes, justifiable overrides, provider non-adherence, and unintended adverse consequences of clinical decision support alerts. The framework was validated by applying it to a medication alerting system for patients with acute kidney injury (AKI). RESULTS: Through expert review, the framework assesses each alert episode for appropriateness of the alert display and the necessity and urgency of a clinical response. Primary outcomes of the framework include the false positive alert rate, alert override rate, provider non-adherence rate, and rate of provider response appropriateness. Application of the framework to evaluate an existing AKI medication alerting system provided a more complete understanding of the process outcomes measured in the AKI medication alerting system. The authors confirmed that previous alerts and provider responses were most often appropriate. CONCLUSION: The new evaluation model offers a potentially effective method for assessing the clinical appropriateness of synchronous interruptive medication alerts prior to evaluating patient outcomes in a comparative trial. More work can determine the generalizability of the framework for use in other settings and other alert types.
Allison B. McCoy, Lemuel R. Waitman, Julia B. Lewis, Julie A. Wright, David P. Choma, Randolph A. Miller, Josh F. Peterson
J. Am. Medical Informatics Assoc.7
2010 Extracting timing and status descriptors for colonoscopy testing from electronic medical records
abstract
Colorectal cancer (CRC) screening rates are low despite confirmed benefits. The authors investigated the use of natural language processing (NLP) to identify previous colonoscopy screening in electronic records from a random sample of 200 patients at least 50 years old. The authors developed algorithms to recognize temporal expressions and 'status indicators', such as 'patient refused', or 'test scheduled'. The new methods were added to the existing KnowledgeMap concept identifier system, and the resulting system was used to parse electronic medical records (EMR) to detect completed colonoscopies. Using as the 'gold standard' expert physicians' manual review of EMR notes, the system identified timing references with a recall of 0.91 and precision of 0.95, colonoscopy status indicators with a recall of 0.82 and precision of 0.95, and references to actually completed colonoscopies with recall of 0.93 and precision of 0.95. The system was superior to using colonoscopy billing codes alone. Health services researchers and clinicians may find NLP a useful adjunct to traditional methods to detect CRC screening status. Further investigations must validate extension of NLP approaches for other types of CRC screening applications.
Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson
J. Am. Medical Informatics Assoc.2
2009 Development of a Natural Language Processing System to Identify Timing and Status of Colonoscopy Testing in Electronic Medical Records
Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson
AMIA2
2009 Research Paper: Evaluation of a Method to Identify and Categorize Section Headers in Clinical Documents
abstract
OBJECTIVE: Clinical notes, typically written in natural language, often contain substructure that divides them into sections, such as "History of Present Illness" or "Family Medical History." The authors designed and evaluated an algorithm ("SecTag") to identify both labeled and unlabeled (implied) note section headers in "history and physical examination" documents ("H&P notes"). DESIGN: The SecTag algorithm uses a combination of natural language processing techniques, word variant recognition with spelling correction, terminology-based rules, and naive Bayesian scoring methods to identify note section headers. Eleven physicians evaluated SecTag's performance on 319 randomly chosen H&P notes. MEASUREMENTS: The primary outcomes were the algorithm's recall and precision in identifying all document sections and a predefined list of twenty-nine major sections. A secondary outcome was to evaluate the algorithm's ability to recognize the correct start and end boundaries of identified sections. RESULTS: The SecTag algorithm identified 16,036 total sections and 7,858 major sections. Physician evaluators classified 15,329 as true positives and identified 160 sections omitted by SecTag. The recall and precision of the SecTag algorithm were 99.0 and 95.6% for all sections, 98.6 and 96.2% for major sections, and 96.6 and 86.8% for unlabeled sections. The algorithm determined the correct starting and ending text boundaries for 94.8% of labeled sections and 85.9% of unlabeled sections. CONCLUSIONS: The SecTag algorithm accurately identified both labeled and unlabeled sections in history and physical documents. This type of algorithm may assist in natural language processing applications, such as clinical decision support systems or competency assessment for medical trainees.
Joshua C. Denny, Anderson Spickard III, Kevin B. Johnson, Neeraja B. Peterson, Josh F. Peterson, Randolph A. Miller
J. Am. Medical Informatics Assoc.5
2007 Research Paper: Medication Administration Discrepancies Persist Despite Electronic Ordering
abstract
Background Up to 38% of inpatient medication errors occur at the administration stage. Although they reduce prescribing errors, computerized provider order entry (CPOE) systems do not prevent administration errors or timing discrepancies. This study determined the degree to which CPOE medication orders matched actual dose administration times. METHODS At a 658-bed academic hospital with CPOE but lacking electronic medication administration charting, authors randomly selected adult patients with eligible medication orders from historical 1999-2003 CPOE log files. Retrospective manual chart audits compared expected (from CPOE) and actual timing of medication administrations. Outcomes included: dose omissions, median lag times between ordered and charted administrations, unauthorized doses, wrong dose errors, and the rate of nurses' medication schedule shifting. RESULTS Dose omissions occurred in 756 of 6019 (12.6%) audited administration opportunities; only 313 of the omissions (5.2% of opportunities) were unexplained. Wrong doses and unexpected doses occurred for 0.1% and 0.7% of opportunities, respectively. Median lag from expected first dose to actual charted administration time was 27 minutes (IQR 0-127). Nursing staff shifted from ordered to alternate administration schedules for 10.7% of regularly scheduled recurring medication orders. Chart review identified reasons for dose omissions, delays, and dose shifting. CONCLUSION Inpatient CPOE orders are legible and conveyed electronically to nurses and the pharmacy. Nonetheless, ward-based medication administrations do not consistently occur as ordered. Medication administration discrepancies are likely to persist even after implementing CPOE and bar-coded medication administration unless recommended interventions are made to address issues such as determining the true urgency of medication administration, avoiding overlapping duplicative medication orders, and developing a safe means for shifting dosing schedules.
Fern FitzHenry, Josh F. Peterson, Mark Arrieta, Lemuel R. Waitman, Jonathan S. Schildcrout, Randolph A. Miller
J. Am. Medical Informatics Assoc.2
2005 Identifying UMLS concepts from ECG Impressions using Knowledge Map
Joshua C. Denny, Anderson Spickard III, Randolph A. Miller, Jonathan S. Schildcrout, Dawood Darbar, S. Trent Rosenbloom, Josh F. Peterson
AMIA7
2005 Measuring the Quality of Medication Administration
Fern FitzHenry, Josh F. Peterson, Mark Arrieta, Randolph A. Miller
AMIA2
2003 Adequacy of representation of the National Drug File Reference Terminology Physiologic Effects reference hierarchy for commonly prescribed medications
S. Trent Rosenbloom, Joseph Awad, Theodore Speroff, Peter L. Elkin, Russell L. Rothman, Anderson Spickard III, Josh F. Peterson, Brent A. Bauer, Dietlind Wahner-Roedler, William M. Gregg, Kevin B. Johnson, Jim Jirjis, Mark Erlbaum, John S. Carter, Michael J. Lincoln, Steven H. Brown
AMIA7
2003 Research Paper: Electronically Screening Discharge Summaries for Adverse Medical Events
abstract
OBJECTIVE: Detecting adverse events is pivotal for measuring and improving medical safety, yet current techniques discourage routine screening. The authors hypothesized that discharge summaries would include information on adverse events, and they developed and evaluated an electronic method for screening medical discharge summaries for adverse events. DESIGN: A cohort study including 424 randomly selected admissions to the medical services of an academic medical center was conducted between January and July 2000. The authors developed a computerized screening tool that searched free-text discharge summaries for trigger words representing possible adverse events. MEASUREMENTS: All discharge summaries with a trigger word present underwent chart review by two independent physician reviewers. The presence of adverse events was assessed using structured implicit judgment. A random sample of discharge summaries without trigger words also was reviewed. RESULTS: Fifty-nine percent (251 of 424) of the discharge summaries contained trigger words. Based on discharge summary review, 44.8% (327 of 730) of the alerted trigger words indicated a possible adverse event. After medical record review, the tool detected 131 adverse events. The sensitivity and specificity of the screening tool were 69% and 48%, respectively. The positive predictive value of the tool was 52%. CONCLUSION: Medical discharge summaries contain information regarding adverse events. Electronic screening of discharge summaries for adverse events using keyword searches is feasible but thus far has poor specificity. Nonetheless, computerized clinical narrative screening methods could potentially offer researchers and quality managers a means to routinely detect adverse events.
Harvey J. Murff, Alan J. Forster, Josh F. Peterson, Julie M. Fiskio, Heather L. Heiman, David W. Bates
J. Am. Medical Informatics Assoc.3
2002 Gerios: Recommending Drugs and Dosing for Elderly Patients
Josh F. Peterson, David W. Bates, Jerry Avorn, Gilad J. Kuperman
AMIA1
2002 Poster Abstract: Electronically Screening Discharge Summaries for Adverse Medical Events
abstract
Detecting and preventing adverse medical events (AEs) is essential for improving medical quality. While electronic approaches for detecting and preventing adverse drug events have been developed, AEs, which include the entire range of events and are thus more diverse, have been harder to detect. Prior studies have detected AEs through structured chart reviews. While this approach is effective, it is costly and time consuming. Thus, we developed a computerized discharge abstract screening tool to detect AEs. Our initial sample consisted of 424 randomly selected patients discharged from the medical services of the Brigham and Women's Hospital between January 1 to June 30, 2000. We developed a set of alert signals based on screening criteria used in the Harvard Medical Practice Study 3 that ultimately including 94 trigger words. Individual trigger words were then identified using text-based searches of the hospital course section of electronically stored discharge summaries. Discharge summaries generating an alert were classified as “screened positive discharge summaries“ and were reviewed to determine the context in which the trigger word had been used and whether an AE appeared likely based on the discharge summary. All screened positive discharge summaries underwent chart review by two independent physician reviewers. The presence of an AE was assessed using structured implicit judgement. A random 25% of screened negative discharge summaries were also reviewed. The positive predictive values for the electronic tool was determined by dividing the number of admissions with discharge summary trigger words and an AE by the total number of screened positive discharge summaries. Time spent reviewing discharge abstracts was recorded. Nine hundred and fifty-three alerts were detected, and after adjusting for repeated signals within the same discharge summary, a total of 733 unique alerts were generated in 251/424 (59%) patients. In 131 screened positive discharge summaries the patient had experienced an AEs based on chart review (kappa statistic = 0.77). Sixty AE occurred within the 173 patients without screened positive discharge summaries. The sensitivity and specificity of the screening tool were 69% and 48% respectively. The positive predictive value of the tool was 52%. The most common category of AE detected was adverse drug events representing 52% of the detected events. The time required to review the screened discharge abstracts was 18 hours. By using computerized screening of discharge abstracts we were able to identify AE's in 31% of the patients sampled. The tool performed reasonable well, however removing individual trigger words with low positive predictive values and other improvements could also improve sensitivity. Using an electronic screening tool, we were able to screen 424 charts in 18 hours. Using Harvard Medical Practice methodology this same initial sample would have required approximately 70 hours. Electronically screening discharge summaries for adverse events appears to be an efficient and feasible means of detecting AE within hospitalized medical patients. Reprinted from the Proceedings of the 2001 AMIA Annual Symposium, with permission.
Harvey J. Murff, Alan J. Forster, Josh F. Peterson, Julie M. Fiskio, Heather L. Heiman, David W. Bates
J. Am. Medical Informatics Assoc.3
2002 Poster Abstract: Drug-Lab Triggers Have Potential to Prevent Adverse Drug Events in Outpatients
abstract
Previous studies have found that many adverse drug events (ADEs) in inpatients can be detected or prevented by alerting physicians to measured physiologic parameters such as an elevated creatinine or hyperkalemia. 1 , 2 In developing a decision support system for an outpatient Electronic Medical Record, we have begun to retrospectively study associations between drugs and labs that could trigger an alert to physicians. It is unknown whether such a drug-lab monitoring system is useful in identifying or preventing ADEs in outpatients. We developed a list of drug-lab triggers using previously published lists and established contraindications for specific drugs. Lab and dose criteria were set using published data when possible and expert opinion when the data was unavailable. For each trigger, we searched the electronic medical records of outpatient clinics in the Partners Health System between 1/1998 and 2/2001. Only the records of patients who were prescribed the drug of interest and had a documented lab of interest were eligible. For each patient record, we determined if the trigger conditions were met at least once during the study period. The proportion of eligible patients who satisfied the trigger criteria at least once is reported in Table 1 . Observed frequencies were calculated by dividing the number of patients who fulfilled all lab and dose criteria by the total number of patients prescribed the drug and had the relevant lab recorded. The most frequent association was between a high allopurinol dose and renal insufficiency occurring in 8.9% of patients prescribed allopurinol. Drug-lab Triggers and Observed Frequency in Outpatients Drug-lab Triggers and Observed Frequency in Outpatients Drug-lab triggers have potential to alert physicians to impending and actual adverse drug events. Events will need to be reviewed in order to calculate a positive predictive value for each trigger. Additionally, confirmation of the utility of triggers will require prospective study of patient outcomes associated with a positive trigger. Reprinted from the Proceedings of the 2001 AMIA Annual Symposium, with permission.
Josh F. Peterson, Deborah H. Williams, Andrew C. Seger, Tejal K. Gandhi, David W. Bates
J. Am. Medical Informatics Assoc.1
2001 Drug-Lab Triggers Have Potential to Prevent Adverse Drug Events in Outpatients
Josh F. Peterson, Deborah H. Williams, Andy Seger, Tejal K. Gandhi, David W. Bates
AMIA1