Bryan D. Steitz

dblp:200/4405 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0003-2066-8692ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 11 first-author · 12 since 2021
YearPublicationVenuePosition
2025 Detecting emergencies in patient portal messages using large language models and knowledge graph-based retrieval-augmented generation
abstract
OBJECTIVES: This study aims to develop and evaluate an approach using large language models (LLMs) and a knowledge graph to triage patient messages that need emergency care. The goal is to notify patients when their messages indicate an emergency, guiding them to seek immediate help rather than using the patient portal, to improve patient safety. MATERIALS AND METHODS: We selected 1020 messages sent to Vanderbilt University Medical Center providers between January 1, 2022 and March 7, 2023. We developed four models to triage these messages for emergencies: (1) Prompt-Only: the patient message was input with a prompt directly into the LLM; (2) Naïve Retrieval Augmented Generation (RAG): provided retrieved information as context to the LLM; (3) RAG from Knowledge Graph with Local Search: a knowledge graph was used to retrieve locally relevant information based on semantic similarities; (4) RAG from Knowledge Graph with Global Search: a knowledge graph was used to retrieve globally relevant information through hierarchical community detection. The knowledge base was a triage book covering 225 protocols. RESULTS: The RAG from Knowledge Graph model with global search outperformed other models, achieving an accuracy of 0.99, a sensitivity of 0.98, and a specificity of 0.99. It demonstrated significant improvements in triaging emergency messages compared to LLM without RAG and naïve RAG. DISCUSSION: The traditional LLM without any retrieval mechanism underperformed compared to models with RAG, which aligns with the expected benefits of augmenting LLMs with domain-specific knowledge sources. Our results suggest that providing external knowledge, especially in a structured manner and in community summaries, can improve LLM performance in triaging patient portal messages. CONCLUSION: LLMs can effectively assist in triaging emergency patient messages after integrating with a knowledge graph about a nurse triage book. Future research should focus on expanding the knowledge graph and deploying the system to evaluate its impact on patient outcomes.
Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.5
2024 Leveraging large language models for generating responses to patient messages - a subjective analysis
abstract
OBJECTIVE: This study aimed to develop and assess the performance of fine-tuned large language models for generating responses to patient messages sent via an electronic health record patient portal. MATERIALS AND METHODS: Utilizing a dataset of messages and responses extracted from the patient portal at a large academic medical center, we developed a model (CLAIR-Short) based on a pre-trained large language model (LLaMA-65B). In addition, we used the OpenAI API to update physician responses from an open-source dataset into a format with informative paragraphs that offered patient education while emphasizing empathy and professionalism. By combining with this dataset, we further fine-tuned our model (CLAIR-Long). To evaluate fine-tuned models, we used 10 representative patient portal questions in primary care to generate responses. We asked primary care physicians to review generated responses from our models and ChatGPT and rated them for empathy, responsiveness, accuracy, and usefulness. RESULTS: The dataset consisted of 499 794 pairs of patient messages and corresponding responses from the patient portal, with 5000 patient messages and ChatGPT-updated responses from an online platform. Four primary care physicians participated in the survey. CLAIR-Short exhibited the ability to generate concise responses similar to provider's responses. CLAIR-Long responses provided increased patient educational content compared to CLAIR-Short and were rated similarly to ChatGPT's responses, receiving positive evaluations for responsiveness, empathy, and accuracy, while receiving a neutral rating for usefulness. CONCLUSION: This subjective analysis suggests that leveraging large language models to generate responses to patient messages demonstrates significant potential in facilitating communication between patients and healthcare providers.
Siru Liu, Allison B. McCoy, Aileen P. Wright, Babatunde Carew, Julian Z. Genkins, Sean S. Huang, Josh F. Peterson, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.8
2024 Using large language model to guide patients to create efficient and comprehensive clinical care message
abstract
OBJECTIVE: This study aims to investigate the feasibility of using Large Language Models (LLMs) to engage with patients at the time they are drafting a question to their healthcare providers, and generate pertinent follow-up questions that the patient can answer before sending their message, with the goal of ensuring that their healthcare provider receives all the information they need to safely and accurately answer the patient's question, eliminating back-and-forth messaging, and the associated delays and frustrations. METHODS: We collected a dataset of patient messages sent between January 1, 2022 to March 7, 2023 at Vanderbilt University Medical Center. Two internal medicine physicians identified 7 common scenarios. We used 3 LLMs to generate follow-up questions: (1) Comprehensive LLM Artificial Intelligence Responder (CLAIR): a locally fine-tuned LLM, (2) GPT4 with a simple prompt, and (3) GPT4 with a complex prompt. Five physicians rated them with the actual follow-ups written by healthcare providers on clarity, completeness, conciseness, and utility. RESULTS: For five scenarios, our CLAIR model had the best performance. The GPT4 model received higher scores for utility and completeness but lower scores for clarity and conciseness. CLAIR generated follow-up questions with similar clarity and conciseness as the actual follow-ups written by healthcare providers, with higher utility than healthcare providers and GPT4, and lower completeness than GPT4, but better than healthcare providers. CONCLUSION: LLMs can generate follow-up patient messages designed to clarify a medical question that compares favorably to those generated by healthcare providers.
Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Julian Z. Genkins, Josh F. Peterson, Yaa A. Kumah-Crystal, William Martinez, Babatunde Carew, Dara Eckerle Mize, Bryan D. Steitz, Adam Wright
J. Am. Medical Informatics Assoc.11
2023 Impact of notification policy on patient-before-clinician review of immediately released test results
abstract
The 21st Century Cures Act mandates immediate availability of test results upon request. The Cures Act does not require that patients be informed of results, but many organizations send notifications when results become available. Our medical center implemented 2 sequential policies: immediate notifications for all results, and notifications only to patients who opt in. We used over 2 years of data from Vanderbilt University Medical Center to measure the effect of these policies on rates of patient-before-clinician result review and patient-initiated messaging using interrupted time series analysis. When releasing test results with immediate notification, the proportion of patient-before-clinician review increased 4-fold and the proportion of patients who sent messages rose 3%. After transition to opt-in notifications, patient-before-clinician review decreased 2.4% and patient-initiated messaging decreased 0.4%. Replacing automated notifications with an opt-in policy provides patients flexibility to indicate their preferences but may not substantially alleviate clinicians' messaging workload.
Bryan D. Steitz, Nana Addo Padi-Adjirackor, Kevin N. Griffith, Thomas J. Reese, S. Trent Rosenbloom, Jessica S. Ancker
J. Am. Medical Informatics Assoc.1
2022 Leveraging Natural Language Processing Tool to Identify Eligible Lung Cancer Screening Patients in the Electronic Health Record
Siru Liu, Allison B. McCoy, Bryan D. Steitz, Adam Wright
AMIA3
2022 Improving the Performance of Large Language Models by Domain-Specific Pre-Training on Clinical Documents
Bryan D. Steitz, Charreau Bell, Jesse Spencer-Smith, Adam Wright
AMIA1
2022 New onset delirium prediction using machine learning and long short-term memory (LSTM) in electronic health record
abstract
OBJECTIVE: To develop and test an accurate deep learning model for predicting new onset delirium in hospitalized adult patients. METHODS: Using electronic health record (EHR) data extracted from a large academic medical center, we developed a model combining long short-term memory (LSTM) and machine learning to predict new onset delirium and compared its performance with machine-learning-only models (logistic regression, random forest, support vector machine, neural network, and LightGBM). The labels of models were confusion assessment method (CAM) assessments. We evaluated models on a hold-out dataset. We calculated Shapley additive explanations (SHAP) measures to gauge the feature impact on the model. RESULTS: A total of 331 489 CAM assessments with 896 features from 34 035 patients were included. The LightGBM model achieved the best performance (AUC 0.927 [0.924, 0.929] and F1 0.626 [0.618, 0.634]) among the machine learning models. When combined with the LSTM model, the final model's performance improved significantly (P = .001) with AUC 0.952 [0.950, 0.955] and F1 0.759 [0.755, 0.765]. The precision value of the combined model improved from 0.497 to 0.751 with a fixed recall of 0.8. Using the mean absolute SHAP values, we identified the top 20 features, including age, heart rate, Richmond Agitation-Sedation Scale score, Morse fall risk score, pulse, respiratory rate, and level of care. CONCLUSION: Leveraging LSTM to capture temporal trends and combining it with the LightGBM model can significantly improve the prediction of new onset delirium, providing an algorithmic basis for the subsequent development of clinical decision support tools for proactive delirium interventions.
Siru Liu, Joseph J. Schlesinger, Allison B. McCoy, Thomas J. Reese, Bryan D. Steitz, Elise M. Russo, Brian Koh, Adam Wright
J. Am. Medical Informatics Assoc.5
2022 Conceptualizing clinical decision support as complex interventions: a meta-analysis of comparative effectiveness trials
abstract
OBJECTIVES: Complex interventions with multiple components and behavior change strategies are increasingly implemented as a form of clinical decision support (CDS) using native electronic health record functionality. Objectives of this study were, therefore, to (1) identify the proportion of randomized controlled trials with CDS interventions that were complex, (2) describe common gaps in the reporting of complexity in CDS research, and (3) determine the impact of increased complexity on CDS effectiveness. MATERIALS AND METHODS: To assess CDS complexity and identify reporting gaps for characterizing CDS interventions, we used the Preferred Reporting Items for Systematic Reviews and Meta-Analyses reporting tool for complex interventions. We evaluated the effect of increased complexity using random-effects meta-analysis. RESULTS: Most included studies evaluated a complex CDS intervention (76%). No studies described use of analytical frameworks or causal pathways. Two studies discussed use of theory but only one fully described the rationale and put it in context of a behavior change. A small but positive effect (standardized mean difference, 0.147; 95% CI, 0.039-0.255; P < .01) in favor of increasing intervention complexity was observed. DISCUSSION: While most CDS studies should classify interventions as complex, opportunities persist for documenting and providing resources in a manner that would enable CDS interventions to be replicated and adapted. Unless reporting of the design, implementation, and evaluation of CDS interventions improves, only slight benefits can be expected. CONCLUSION: Conceptualizing CDS as complex interventions may help convey the careful attention that is needed to ensure these interventions are contextually and theoretically informed.
Thomas J. Reese, Siru Liu, Bryan D. Steitz, Allison B. McCoy, Elise M. Russo, Brian Koh, Jessica S. Ancker, Adam Wright
J. Am. Medical Informatics Assoc.3
2021 Evaluating the Scope of Collaboration Among Primary Care Teams through Electronic Asynchronous Communication
Arianna E. Nimocks, Bryan D. Steitz, Adam Wright
AMIA2
2021 Applying State of the Art Language Models to Enable Better Clinical Natural Language Processing
Bryan D. Steitz, Emily Alsentzer, Hoo Chang Shin, Byron C. Wallace, Adam Wright
AMIA1
2021 Evaluating Primary Care Provider Work of Managing Asynchronous Messages through Electronic Health Record Access Logs
Bryan D. Steitz, Adam Wright
AMIA1
2021 Comparing Language Model Vocabulary Coverage on Clinical Documents
Bryan D. Steitz, Adam Wright
AMIA1
2020 Quantifying the Hidden Electronic Work of Clinical Communication in a Breast Cancer Cohort
Bryan D. Steitz, Kim M. Unertl, Mia A. Levy
AMIA1
2020 Characterizing communication patterns among members of the clinical care team to deliver breast cancer treatment
abstract
OBJECTIVE: Research to date focused on quantifying team collaboration has relied on identifying shared patients but does not incorporate the major role of communication patterns. The goal of this study was to describe the patterns and volume of communication among care team members involved in treating breast cancer patients. MATERIALS AND METHODS: We analyzed 4 years of communications data from the electronic health record between care team members at Vanderbilt University Medical Center (VUMC). Our cohort of patients diagnosed with breast cancer was identified using the VUMC tumor registry. We classified each care team member participating in electronic messaging by their institutional role and classified physicians by specialty. To identify collaborative patterns, we modeled the data as a social network. RESULTS: Our cohort of 1181 patients was the subject of 322 424 messages sent in 104 210 unique communication threads by 5620 employees. On average, each patient was the subject of 88.2 message threads involving 106.4 employees. Each employee, on average, sent 72.9 messages and was connected to 24.6 collaborators. Nurses and physicians were involved in 98% and 44% of all message threads, respectively. DISCUSSION AND CONCLUSION: Our results suggest that many providers in our study may experience a high volume of messaging work. By using data routinely generated through interaction with the electronic health record, we can begin to evaluate how to iteratively implement and assess initiatives to improve the efficiency of care coordination and reduce unnecessary messaging work across all care team roles.
Bryan D. Steitz, Kim M. Unertl, Mia A. Levy
J. Am. Medical Informatics Assoc.1
2018 Patient Experience During Electronic Health Record Migration
Bryan D. Steitz, Brian Carlson, Kim M. Unertl
AMIA1
2016 A Social Network Analysis of Cancer Provider Collaboration
Bryan D. Steitz, Mia A. Levy
AMIA1
2015 Bringing Context to Data Analytics: A Hybrid Approach to Understanding Clinical Workflow
Bryan D. Steitz, Kim M. Unertl
AMIA1
2014 Collaborative mHealth Tools for Diabetes Management
Bryan D. Steitz, Erika S. Poole, Madhu C. Reddy
AMIA1