Aileen P. Wright

dblp:161/1909 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-7550-4284ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Enterprise-wide simultaneous deployment of ambient scribe technology: lessons learned from an academic health system
abstract
OBJECTIVES: To report on the feasibility of a simultaneous, enterprise-wide deployment of EHR-integrated ambient scribe technology across a large academic health system. MATERIALS AND METHODS: On January 15, 2025, ambient scribing was made available to over 2400 ambulatory and emergency department clinicians. We tracked utilization rates, technical support needs, and user feedback. RESULTS: By March 31, 2025, 20.1% of visit notes incorporated ambient scribing, and 1223 clinicians had used ambient scribing. Among 209 respondents (22.1% of 947 surveyed), 90.9% would be disappointed if they lost access to ambient scribing, and 84.7% reported a positive training experience. DISCUSSION: Enterprise-wide simultaneous deployment combined with a low-barrier training model enabled immediate access for clinicians and reduced administrative burden by concentrating go-live efforts. Support needs were manageable. CONCLUSION: Simultaneous enterprise-wide deployment of ambient scribing was feasible and provided immediate access for clinicians.
Aileen P. Wright, Carolynn K. Nall, Jacob Franklin, Sara N. Horst, Yaa A. Kumah-Crystal, Adam Wright, Dara Eckerle Mize
J. Am. Medical Informatics Assoc.1
2025 Detecting emergencies in patient portal messages using large language models and knowledge graph-based retrieval-augmented generation
abstract
OBJECTIVES: This study aims to develop and evaluate an approach using large language models (LLMs) and a knowledge graph to triage patient messages that need emergency care. The goal is to notify patients when their messages indicate an emergency, guiding them to seek immediate help rather than using the patient portal, to improve patient safety. MATERIALS AND METHODS: We selected 1020 messages sent to Vanderbilt University Medical Center providers between January 1, 2022 and March 7, 2023. We developed four models to triage these messages for emergencies: (1) Prompt-Only: the patient message was input with a prompt directly into the LLM; (2) Naïve Retrieval Augmented Generation (RAG): provided retrieved information as context to the LLM; (3) RAG from Knowledge Graph with Local Search: a knowledge graph was used to retrieve locally relevant information based on semantic similarities; (4) RAG from Knowledge Graph with Global Search: a knowledge graph was used to retrieve globally relevant information through hierarchical community detection. The knowledge base was a triage book covering 225 protocols. RESULTS: The RAG from Knowledge Graph model with global search outperformed other models, achieving an accuracy of 0.99, a sensitivity of 0.98, and a specificity of 0.99. It demonstrated significant improvements in triaging emergency messages compared to LLM without RAG and naïve RAG. DISCUSSION: The traditional LLM without any retrieval mechanism underperformed compared to models with RAG, which aligns with the expected benefits of augmenting LLMs with domain-specific knowledge sources. Our results suggest that providing external knowledge, especially in a structured manner and in community summaries, can improve LLM performance in triaging patient portal messages. CONCLUSION: LLMs can effectively assist in triaging emergency patient messages after integrating with a knowledge graph about a nurse triage book. Future research should focus on expanding the knowledge graph and deploying the system to evaluate its impact on patient outcomes.
Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.2
2024 Leveraging large language models for generating responses to patient messages - a subjective analysis
abstract
OBJECTIVE: This study aimed to develop and assess the performance of fine-tuned large language models for generating responses to patient messages sent via an electronic health record patient portal. MATERIALS AND METHODS: Utilizing a dataset of messages and responses extracted from the patient portal at a large academic medical center, we developed a model (CLAIR-Short) based on a pre-trained large language model (LLaMA-65B). In addition, we used the OpenAI API to update physician responses from an open-source dataset into a format with informative paragraphs that offered patient education while emphasizing empathy and professionalism. By combining with this dataset, we further fine-tuned our model (CLAIR-Long). To evaluate fine-tuned models, we used 10 representative patient portal questions in primary care to generate responses. We asked primary care physicians to review generated responses from our models and ChatGPT and rated them for empathy, responsiveness, accuracy, and usefulness. RESULTS: The dataset consisted of 499 794 pairs of patient messages and corresponding responses from the patient portal, with 5000 patient messages and ChatGPT-updated responses from an online platform. Four primary care physicians participated in the survey. CLAIR-Short exhibited the ability to generate concise responses similar to provider's responses. CLAIR-Long responses provided increased patient educational content compared to CLAIR-Short and were rated similarly to ChatGPT's responses, receiving positive evaluations for responsiveness, empathy, and accuracy, while receiving a neutral rating for usefulness. CONCLUSION: This subjective analysis suggests that leveraging large language models to generate responses to patient messages demonstrates significant potential in facilitating communication between patients and healthcare providers.
Siru Liu, Allison B. McCoy, Aileen P. Wright, Babatunde Carew, Julian Z. Genkins, Sean S. Huang, Josh F. Peterson, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.3
2024 Why do users override alerts? Utilizing large language model to summarize comments and optimize clinical decision support
abstract
OBJECTIVES: To evaluate the capability of using generative artificial intelligence (AI) in summarizing alert comments and to determine if the AI-generated summary could be used to improve clinical decision support (CDS) alerts. MATERIALS AND METHODS: We extracted user comments to alerts generated from September 1, 2022 to September 1, 2023 at Vanderbilt University Medical Center. For a subset of 8 alerts, comment summaries were generated independently by 2 physicians and then separately by GPT-4. We surveyed 5 CDS experts to rate the human-generated and AI-generated summaries on a scale from 1 (strongly disagree) to 5 (strongly agree) for the 4 metrics: clarity, completeness, accuracy, and usefulness. RESULTS: Five CDS experts participated in the survey. A total of 16 human-generated summaries and 8 AI-generated summaries were assessed. Among the top 8 rated summaries, five were generated by GPT-4. AI-generated summaries demonstrated high levels of clarity, accuracy, and usefulness, similar to the human-generated summaries. Moreover, AI-generated summaries exhibited significantly higher completeness and usefulness compared to the human-generated summaries (AI: 3.4 ± 1.2, human: 2.7 ± 1.2, P = .001). CONCLUSION: End-user comments provide clinicians' immediate feedback to CDS alerts and can serve as a direct and valuable data resource for improving CDS delivery. Traditionally, these comments may not be considered in the CDS review process due to their unstructured nature, large volume, and the presence of redundant or irrelevant content. Our study demonstrates that GPT-4 is capable of distilling these comments into summaries characterized by high clarity, accuracy, and completeness. AI-generated summaries are equivalent and potentially better than human-generated summaries. These AI-generated summaries could provide CDS experts with a novel means of reviewing user comments to rapidly optimize CDS alerts both online and offline.
Siru Liu, Allison B. McCoy, Aileen P. Wright, Scott D. Nelson, Sean S. Huang, Hasan B. Ahmad, Sabrina E. Carro, Jacob Franklin, James Brogan, Adam Wright
J. Am. Medical Informatics Assoc.3
2024 Using large language model to guide patients to create efficient and comprehensive clinical care message
abstract
OBJECTIVE: This study aims to investigate the feasibility of using Large Language Models (LLMs) to engage with patients at the time they are drafting a question to their healthcare providers, and generate pertinent follow-up questions that the patient can answer before sending their message, with the goal of ensuring that their healthcare provider receives all the information they need to safely and accurately answer the patient's question, eliminating back-and-forth messaging, and the associated delays and frustrations. METHODS: We collected a dataset of patient messages sent between January 1, 2022 to March 7, 2023 at Vanderbilt University Medical Center. Two internal medicine physicians identified 7 common scenarios. We used 3 LLMs to generate follow-up questions: (1) Comprehensive LLM Artificial Intelligence Responder (CLAIR): a locally fine-tuned LLM, (2) GPT4 with a simple prompt, and (3) GPT4 with a complex prompt. Five physicians rated them with the actual follow-ups written by healthcare providers on clarity, completeness, conciseness, and utility. RESULTS: For five scenarios, our CLAIR model had the best performance. The GPT4 model received higher scores for utility and completeness but lower scores for clarity and conciseness. CLAIR generated follow-up questions with similar clarity and conciseness as the actual follow-ups written by healthcare providers, with higher utility than healthcare providers and GPT4, and lower completeness than GPT4, but better than healthcare providers. CONCLUSION: LLMs can generate follow-up patient messages designed to clarify a medical question that compares favorably to those generated by healthcare providers.
Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Julian Z. Genkins, Josh F. Peterson, Yaa A. Kumah-Crystal, William Martinez, Babatunde Carew, Dara Eckerle Mize, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.2
2023 Using AI-generated suggestions from ChatGPT to optimize clinical decision support
abstract
OBJECTIVE: To determine if ChatGPT can generate useful suggestions for improving clinical decision support (CDS) logic and to assess noninferiority compared to human-generated suggestions. METHODS: We supplied summaries of CDS logic to ChatGPT, an artificial intelligence (AI) tool for question answering that uses a large language model, and asked it to generate suggestions. We asked human clinician reviewers to review the AI-generated suggestions as well as human-generated suggestions for improving the same CDS alerts, and rate the suggestions for their usefulness, acceptance, relevance, understanding, workflow, bias, inversion, and redundancy. RESULTS: Five clinicians analyzed 36 AI-generated suggestions and 29 human-generated suggestions for 7 alerts. Of the 20 suggestions that scored highest in the survey, 9 were generated by ChatGPT. The suggestions generated by AI were found to offer unique perspectives and were evaluated as highly understandable and relevant, with moderate usefulness, low acceptance, bias, inversion, redundancy. CONCLUSION: AI-generated suggestions could be an important complementary part of optimizing CDS alerts, can identify potential improvements to alert logic and support their implementation, and may even be able to assist experts in formulating their own suggestions for CDS improvement. ChatGPT shows great potential for using large language models and reinforcement learning from human feedback to improve CDS alert logic and potentially other medical areas involving complex, clinical logic, a key step in the development of an advanced learning health system.
Siru Liu, Aileen P. Wright, Barron L. Patterson, Jonathan P. Wanderer, Robert W. Turer, Scott D. Nelson, Allison B. McCoy, Dean F. Sittig, Adam Wright
J. Am. Medical Informatics Assoc.2
2022 Clinician collaboration to improve clinical decision support: the Clickbusters initiative
abstract
OBJECTIVE: We describe the Clickbusters initiative implemented at Vanderbilt University Medical Center (VUMC), which was designed to improve safety and quality and reduce burnout through the optimization of clinical decision support (CDS) alerts. MATERIALS AND METHODS: We developed a 10-step Clickbusting process and implemented a program that included a curriculum, CDS alert inventory, oversight process, and gamification. We carried out two 3-month rounds of the Clickbusters program at VUMC. We completed descriptive analyses of the changes made to alerts during the process, and of alert firing rates before and after the program. RESULTS: Prior to Clickbusters, VUMC had 419 CDS alerts in production, with 488 425 firings (42 982 interruptive) each week. After 2 rounds, the Clickbusters program resulted in detailed, comprehensive reviews of 84 CDS alerts and reduced the number of weekly alert firings by more than 70 000 (15.43%). In addition to the direct improvements in CDS, the initiative also increased user engagement and involvement in CDS. CONCLUSIONS: At VUMC, the Clickbusters program was successful in optimizing CDS alerts by reducing alert firings and resulting clicks. The program also involved more users in the process of evaluating and improving CDS and helped build a culture of continuous evaluation and improvement of clinical content in the electronic health record.
Allison B. McCoy, Elise M. Russo, Kevin B. Johnson, Bobby Addison, Neal Patel, Jonathan P. Wanderer, Dara Eckerle Mize, Jon G. Jackson, Thomas J. Reese, Sylinda Littlejohn, Lorraine Patterson, Tina French, Debbie Preston, Audra Rosenbury, Charlie Valdez, Scott D. Nelson, Chetan V. Aher, Mhd Wael Alrifai, Jennifer Andrews, Cheryl M. Cobb, Sara N. Horst, David P. Johnson, Lindsey A. Knake, Adam A. Lewis, Laura Parks, Sharidan K. Parr, Pratik Patel, Barron L. Patterson, Christine M. Smith, Krystle D. Suszter, Robert W. Turer, Lyndy J. Wilcox, Aileen P. Wright, Adam Wright
J. Am. Medical Informatics Assoc.33
2020 The Medication Ordering Safety System: A Pipeline to Predict Adverse Drug Events at the Time of Ordering
Aileen P. Wright, Dara Eckerle Mize
AMIA1
2018 Smashing the strict hierarchy: three cases of clinical decision support malfunctions involving carvedilol
abstract
Clinical vocabularies allow for standard representation of clinical concepts, and can also contain knowledge structures, such as hierarchy, that facilitate the creation of maintainable and accurate clinical decision support (CDS). A key architectural feature of clinical hierarchies is how they handle parent-child relationships - specifically whether hierarchies are strict hierarchies (allowing a single parent per concept) or polyhierarchies (allowing multiple parents per concept). These structures handle subsumption relationships (ie, ancestor and descendant relationships) differently. In this paper, we describe three real-world malfunctions of clinical decision support related to incorrect assumptions about subsumption checking for β-blocker, specifically carvedilol, a non-selective β-blocker that also has α-blocker activity. We recommend that 1) CDS implementers should learn about the limitations of terminologies, hierarchies, and classification, 2) CDS implementers should thoroughly test CDS, with a focus on special or unusual cases, 3) CDS implementers should monitor feedback from users, and 4) electronic health record (EHR) and clinical content developers should offer and support polyhierarchical clinical terminologies, especially for medications.
Adam Wright, Aileen P. Wright, Skye Aaron, Dean F. Sittig
J. Am. Medical Informatics Assoc.2
2015 The use of sequential pattern mining to predict next prescribed medications
Aileen P. Wright, Adam Wright, Allison B. McCoy, Dean F. Sittig
J. Biomed. Informatics1