VLDB 2026 Research / reviewers in the wild / expert
Siru Liu
dblp:134/2752
· DBLP profile ↗
21ranked-venue papers
16as first author
16since 2021 · last 2025
0000-0002-5003-5354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 16 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelinesabstractOBJECTIVE: The objectives of this study are to synthesize findings from recent research of retrieval-augmented generation (RAG) and large language models (LLMs) in biomedicine and provide clinical development guidelines to improve effectiveness. MATERIALS AND METHODS: We conducted a systematic literature review and a meta-analysis. The report was created in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses 2020 analysis. Searches were performed in 3 databases (PubMed, Embase, PsycINFO) using terms related to "retrieval augmented generation" and "large language model," for articles published in 2023 and 2024. We selected studies that compared baseline LLM performance with RAG performance. We developed a random-effect meta-analysis model, using odds ratio as the effect size. RESULTS: Among 335 studies, 20 were included in this literature review. The pooled effect size was 1.35, with a 95% confidence interval of 1.19-1.53, indicating a statistically significant effect (P = .001). We reported clinical tasks, baseline LLMs, retrieval sources and strategies, as well as evaluation methods. DISCUSSION: Building on our literature review, we developed Guidelines for Unified Implementation and Development of Enhanced LLM Applications with RAG in Clinical Settings to inform clinical applications using RAG. CONCLUSION: Overall, RAG implementation showed a 1.35 odds ratio increase in performance compared to baseline LLMs. Future research should focus on (1) system-level enhancement: the combination of RAG and agent, (2) knowledge-level enhancement: deep integration of knowledge into LLM, and (3) integration-level enhancement: integrating RAG systems within electronic health records. Siru Liu, Allison B. McCoy, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | Detecting emergencies in patient portal messages using large language models and knowledge graph-based retrieval-augmented generationabstractOBJECTIVES: This study aims to develop and evaluate an approach using large language models (LLMs) and a knowledge graph to triage patient messages that need emergency care. The goal is to notify patients when their messages indicate an emergency, guiding them to seek immediate help rather than using the patient portal, to improve patient safety. MATERIALS AND METHODS: We selected 1020 messages sent to Vanderbilt University Medical Center providers between January 1, 2022 and March 7, 2023. We developed four models to triage these messages for emergencies: (1) Prompt-Only: the patient message was input with a prompt directly into the LLM; (2) Naïve Retrieval Augmented Generation (RAG): provided retrieved information as context to the LLM; (3) RAG from Knowledge Graph with Local Search: a knowledge graph was used to retrieve locally relevant information based on semantic similarities; (4) RAG from Knowledge Graph with Global Search: a knowledge graph was used to retrieve globally relevant information through hierarchical community detection. The knowledge base was a triage book covering 225 protocols. RESULTS: The RAG from Knowledge Graph model with global search outperformed other models, achieving an accuracy of 0.99, a sensitivity of 0.98, and a specificity of 0.99. It demonstrated significant improvements in triaging emergency messages compared to LLM without RAG and naïve RAG. DISCUSSION: The traditional LLM without any retrieval mechanism underperformed compared to models with RAG, which aligns with the expected benefits of augmenting LLMs with domain-specific knowledge sources. Our results suggest that providing external knowledge, especially in a structured manner and in community summaries, can improve LLM performance in triaging patient portal messages. CONCLUSION: LLMs can effectively assist in triaging emergency patient messages after integrating with a knowledge graph about a nurse triage book. Future research should focus on expanding the knowledge graph and deploying the system to evaluate its impact on patient outcomes. Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Bryan D. Steitz, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Leveraging explainable artificial intelligence to optimize clinical decision supportabstractOBJECTIVE: To develop and evaluate a data-driven process to generate suggestions for improving alert criteria using explainable artificial intelligence (XAI) approaches. METHODS: We extracted data on alerts generated from January 1, 2019 to December 31, 2020, at Vanderbilt University Medical Center. We developed machine learning models to predict user responses to alerts. We applied XAI techniques to generate global explanations and local explanations. We evaluated the generated suggestions by comparing with alert's historical change logs and stakeholder interviews. Suggestions that either matched (or partially matched) changes already made to the alert or were considered clinically correct were classified as helpful. RESULTS: The final dataset included 2 991 823 firings with 2689 features. Among the 5 machine learning models, the LightGBM model achieved the highest Area under the ROC Curve: 0.919 [0.918, 0.920]. We identified 96 helpful suggestions. A total of 278 807 firings (9.3%) could have been eliminated. Some of the suggestions also revealed workflow and education issues. CONCLUSION: We developed a data-driven process to generate suggestions for improving alert criteria using XAI techniques. Our approach could identify improvements regarding clinical decision support (CDS) that might be overlooked or delayed in manual reviews. It also unveils a secondary purpose for the XAI: to improve quality by discovering scenarios where CDS alerts are not accepted due to workflow, education, or staffing issues. Siru Liu, Allison B. McCoy, Josh F. Peterson, Thomas A. Lasko, Dean F. Sittig, Scott D. Nelson, Jennifer Andrews, Lorraine Patterson, Cheryl M. Cobb, David Mulherin, Colleen T. Morton, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Leveraging large language models for generating responses to patient messages - a subjective analysisabstractOBJECTIVE: This study aimed to develop and assess the performance of fine-tuned large language models for generating responses to patient messages sent via an electronic health record patient portal. MATERIALS AND METHODS: Utilizing a dataset of messages and responses extracted from the patient portal at a large academic medical center, we developed a model (CLAIR-Short) based on a pre-trained large language model (LLaMA-65B). In addition, we used the OpenAI API to update physician responses from an open-source dataset into a format with informative paragraphs that offered patient education while emphasizing empathy and professionalism. By combining with this dataset, we further fine-tuned our model (CLAIR-Long). To evaluate fine-tuned models, we used 10 representative patient portal questions in primary care to generate responses. We asked primary care physicians to review generated responses from our models and ChatGPT and rated them for empathy, responsiveness, accuracy, and usefulness. RESULTS: The dataset consisted of 499 794 pairs of patient messages and corresponding responses from the patient portal, with 5000 patient messages and ChatGPT-updated responses from an online platform. Four primary care physicians participated in the survey. CLAIR-Short exhibited the ability to generate concise responses similar to provider's responses. CLAIR-Long responses provided increased patient educational content compared to CLAIR-Short and were rated similarly to ChatGPT's responses, receiving positive evaluations for responsiveness, empathy, and accuracy, while receiving a neutral rating for usefulness. CONCLUSION: This subjective analysis suggests that leveraging large language models to generate responses to patient messages demonstrates significant potential in facilitating communication between patients and healthcare providers. Siru Liu, Allison B. McCoy, Aileen P. Wright, Babatunde Carew, Julian Z. Genkins, Sean S. Huang, Josh F. Peterson, Bryan D. Steitz, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Why do users override alerts? Utilizing large language model to summarize comments and optimize clinical decision supportabstractOBJECTIVES: To evaluate the capability of using generative artificial intelligence (AI) in summarizing alert comments and to determine if the AI-generated summary could be used to improve clinical decision support (CDS) alerts. MATERIALS AND METHODS: We extracted user comments to alerts generated from September 1, 2022 to September 1, 2023 at Vanderbilt University Medical Center. For a subset of 8 alerts, comment summaries were generated independently by 2 physicians and then separately by GPT-4. We surveyed 5 CDS experts to rate the human-generated and AI-generated summaries on a scale from 1 (strongly disagree) to 5 (strongly agree) for the 4 metrics: clarity, completeness, accuracy, and usefulness. RESULTS: Five CDS experts participated in the survey. A total of 16 human-generated summaries and 8 AI-generated summaries were assessed. Among the top 8 rated summaries, five were generated by GPT-4. AI-generated summaries demonstrated high levels of clarity, accuracy, and usefulness, similar to the human-generated summaries. Moreover, AI-generated summaries exhibited significantly higher completeness and usefulness compared to the human-generated summaries (AI: 3.4 ± 1.2, human: 2.7 ± 1.2, P = .001). CONCLUSION: End-user comments provide clinicians' immediate feedback to CDS alerts and can serve as a direct and valuable data resource for improving CDS delivery. Traditionally, these comments may not be considered in the CDS review process due to their unstructured nature, large volume, and the presence of redundant or irrelevant content. Our study demonstrates that GPT-4 is capable of distilling these comments into summaries characterized by high clarity, accuracy, and completeness. AI-generated summaries are equivalent and potentially better than human-generated summaries. These AI-generated summaries could provide CDS experts with a novel means of reviewing user comments to rapidly optimize CDS alerts both online and offline. Siru Liu, Allison B. McCoy, Aileen P. Wright, Scott D. Nelson, Sean S. Huang, Hasan B. Ahmad, Sabrina E. Carro, Jacob Franklin, James Brogan, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Using large language model to guide patients to create efficient and comprehensive clinical care messageabstractOBJECTIVE: This study aims to investigate the feasibility of using Large Language Models (LLMs) to engage with patients at the time they are drafting a question to their healthcare providers, and generate pertinent follow-up questions that the patient can answer before sending their message, with the goal of ensuring that their healthcare provider receives all the information they need to safely and accurately answer the patient's question, eliminating back-and-forth messaging, and the associated delays and frustrations. METHODS: We collected a dataset of patient messages sent between January 1, 2022 to March 7, 2023 at Vanderbilt University Medical Center. Two internal medicine physicians identified 7 common scenarios. We used 3 LLMs to generate follow-up questions: (1) Comprehensive LLM Artificial Intelligence Responder (CLAIR): a locally fine-tuned LLM, (2) GPT4 with a simple prompt, and (3) GPT4 with a complex prompt. Five physicians rated them with the actual follow-ups written by healthcare providers on clarity, completeness, conciseness, and utility. RESULTS: For five scenarios, our CLAIR model had the best performance. The GPT4 model received higher scores for utility and completeness but lower scores for clarity and conciseness. CLAIR generated follow-up questions with similar clarity and conciseness as the actual follow-ups written by healthcare providers, with higher utility than healthcare providers and GPT4, and lower completeness than GPT4, but better than healthcare providers. CONCLUSION: LLMs can generate follow-up patient messages designed to clarify a medical question that compares favorably to those generated by healthcare providers. Siru Liu, Aileen P. Wright, Allison B. McCoy, Sean S. Huang, Julian Z. Genkins, Josh F. Peterson, Yaa A. Kumah-Crystal, William Martinez, Babatunde Carew, Dara Eckerle Mize, Bryan D. Steitz, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Using AI-generated suggestions from ChatGPT to optimize clinical decision supportabstractOBJECTIVE: To determine if ChatGPT can generate useful suggestions for improving clinical decision support (CDS) logic and to assess noninferiority compared to human-generated suggestions. METHODS: We supplied summaries of CDS logic to ChatGPT, an artificial intelligence (AI) tool for question answering that uses a large language model, and asked it to generate suggestions. We asked human clinician reviewers to review the AI-generated suggestions as well as human-generated suggestions for improving the same CDS alerts, and rate the suggestions for their usefulness, acceptance, relevance, understanding, workflow, bias, inversion, and redundancy. RESULTS: Five clinicians analyzed 36 AI-generated suggestions and 29 human-generated suggestions for 7 alerts. Of the 20 suggestions that scored highest in the survey, 9 were generated by ChatGPT. The suggestions generated by AI were found to offer unique perspectives and were evaluated as highly understandable and relevant, with moderate usefulness, low acceptance, bias, inversion, redundancy. CONCLUSION: AI-generated suggestions could be an important complementary part of optimizing CDS alerts, can identify potential improvements to alert logic and support their implementation, and may even be able to assist experts in formulating their own suggestions for CDS improvement. ChatGPT shows great potential for using large language models and reinforcement learning from human feedback to improve CDS alert logic and potentially other medical areas involving complex, clinical logic, a key step in the development of an advanced learning health system. Siru Liu, Aileen P. Wright, Barron L. Patterson, Jonathan P. Wanderer, Robert W. Turer, Scott D. Nelson, Allison B. McCoy, Dean F. Sittig, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Part Time MD Program Students' Perception Towards Clinical Informatics
Siru Liu |
AMIA | 2 |
| 2022 | Leveraging Natural Language Processing Tool to Identify Eligible Lung Cancer Screening Patients in the Electronic Health Record
Siru Liu, Allison B. McCoy, Bryan D. Steitz, Adam Wright |
AMIA | 1 |
| 2022 | A Theory-based Evaluation of a Clinical Decision Support System to Predict New Onset of Delirium
Siru Liu, Adam Wright, Joseph J. Schlesinger, Thomas J. Reese, Edward T. Qian, Elise M. Russo, Matthew W. Semler, Brian J. Douthit, Allison B. McCoy |
AMIA | 1 |
| 2022 | An interpretable DIC risk prediction model based on convolutional neural networks with time series dataabstractDisseminated intravascular coagulation (DIC) is a complex, life-threatening syndrome associated with the end-stage of different coagulation disorders. Early prediction of the risk of DIC development is an urgent clinical need to reduce adverse outcomes. However, effective approaches and models to identify early DIC are still lacking. In this study, a novel interpretable deep learning based time series is used to predict the risk of DIC. The study cohort included ICU patients from a 4300-bed academic hospital between January 1, 2019, and January 1, 2022. Experimental results show that our model achieves excellent performance (AUC: 0.986, Accuracy: 95.7%, and F1:0.935). Gradient-weighted Class Activation Mapping (Grad-CAM) was used to explain how predictive models identified patients with DIC. The decision basis of the model was displayed in the form of a heat map. The model can be used to identify high-risk patients with DIC early, which will help in the early intervention of DIC patients and improve the treatment effect. Siru Liu |
BMC Bioinform. | 3 |
| 2022 | The potential for leveraging machine learning to filter medication alertsabstractOBJECTIVE: To evaluate the potential for machine learning to predict medication alerts that might be ignored by a user, and intelligently filter out those alerts from the user's view. MATERIALS AND METHODS: We identified features (eg, patient and provider characteristics) proposed to modulate user responses to medication alerts through the literature; these features were then refined through expert review. Models were developed using rule-based and machine learning techniques (logistic regression, random forest, support vector machine, neural network, and LightGBM). We collected log data on alerts shown to users throughout 2019 at University of Utah Health. We sought to maximize precision while maintaining a false-negative rate <0.01, a threshold predefined through discussion with physicians and pharmacists. We developed models while maintaining a sensitivity of 0.99. Two null hypotheses were developed: H1-there is no difference in precision among prediction models; and H2-the removal of any feature category does not change precision. RESULTS: A total of 3,481,634 medication alerts with 751 features were evaluated. With sensitivity fixed at 0.99, LightGBM achieved the highest precision of 0.192 and less than 0.01 for the pre-defined maximal false-negative rate by subject-matter experts (H1) (P < 0.001). This model could reduce alert volume by 54.1%. We removed different combinations of features (H2) and found that not all features significantly contributed to precision. Removing medication order features (eg, dosage) most significantly decreased precision (-0.147, P = 0.001). CONCLUSIONS: Machine learning potentially enables the intelligent filtering of medication alerts. Siru Liu, Kensaku Kawamoto, Guilherme Del Fiol, Charlene R. Weir, Daniel C. Malone, Thomas J. Reese, Keaton L. Morgan, David El Halta, Samir E. AbdelRahman |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | New onset delirium prediction using machine learning and long short-term memory (LSTM) in electronic health recordabstractOBJECTIVE: To develop and test an accurate deep learning model for predicting new onset delirium in hospitalized adult patients. METHODS: Using electronic health record (EHR) data extracted from a large academic medical center, we developed a model combining long short-term memory (LSTM) and machine learning to predict new onset delirium and compared its performance with machine-learning-only models (logistic regression, random forest, support vector machine, neural network, and LightGBM). The labels of models were confusion assessment method (CAM) assessments. We evaluated models on a hold-out dataset. We calculated Shapley additive explanations (SHAP) measures to gauge the feature impact on the model. RESULTS: A total of 331 489 CAM assessments with 896 features from 34 035 patients were included. The LightGBM model achieved the best performance (AUC 0.927 [0.924, 0.929] and F1 0.626 [0.618, 0.634]) among the machine learning models. When combined with the LSTM model, the final model's performance improved significantly (P = .001) with AUC 0.952 [0.950, 0.955] and F1 0.759 [0.755, 0.765]. The precision value of the combined model improved from 0.497 to 0.751 with a fixed recall of 0.8. Using the mean absolute SHAP values, we identified the top 20 features, including age, heart rate, Richmond Agitation-Sedation Scale score, Morse fall risk score, pulse, respiratory rate, and level of care. CONCLUSION: Leveraging LSTM to capture temporal trends and combining it with the LightGBM model can significantly improve the prediction of new onset delirium, providing an algorithmic basis for the subsequent development of clinical decision support tools for proactive delirium interventions. Siru Liu, Joseph J. Schlesinger, Allison B. McCoy, Thomas J. Reese, Bryan D. Steitz, Elise M. Russo, Brian Koh, Adam Wright |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Conceptualizing clinical decision support as complex interventions: a meta-analysis of comparative effectiveness trialsabstractOBJECTIVES: Complex interventions with multiple components and behavior change strategies are increasingly implemented as a form of clinical decision support (CDS) using native electronic health record functionality. Objectives of this study were, therefore, to (1) identify the proportion of randomized controlled trials with CDS interventions that were complex, (2) describe common gaps in the reporting of complexity in CDS research, and (3) determine the impact of increased complexity on CDS effectiveness. MATERIALS AND METHODS: To assess CDS complexity and identify reporting gaps for characterizing CDS interventions, we used the Preferred Reporting Items for Systematic Reviews and Meta-Analyses reporting tool for complex interventions. We evaluated the effect of increased complexity using random-effects meta-analysis. RESULTS: Most included studies evaluated a complex CDS intervention (76%). No studies described use of analytical frameworks or causal pathways. Two studies discussed use of theory but only one fully described the rationale and put it in context of a behavior change. A small but positive effect (standardized mean difference, 0.147; 95% CI, 0.039-0.255; P < .01) in favor of increasing intervention complexity was observed. DISCUSSION: While most CDS studies should classify interventions as complex, opportunities persist for documenting and providing resources in a manner that would enable CDS interventions to be replicated and adapted. Unless reporting of the design, implementation, and evaluation of CDS interventions improves, only slight benefits can be expected. CONCLUSION: Conceptualizing CDS as complex interventions may help convey the careful attention that is needed to ensure these interventions are contextually and theoretically informed. Thomas J. Reese, Siru Liu, Bryan D. Steitz, Allison B. McCoy, Elise M. Russo, Brian Koh, Jessica S. Ancker, Adam Wright |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Leveraging Transfer Learning to Analyze Opinions, Attitudes, and Behavioral Intentions Toward COVID-19 Vaccines
Siru Liu |
AMIA | 1 |
| 2021 | A theory-based meta-regression of factors influencing clinical decision support adoption and implementationabstractOBJECTIVE: The purpose of the study was to explore the theoretical underpinnings of effective clinical decision support (CDS) factors using the comparative effectiveness results. MATERIALS AND METHODS: We leveraged search results from a previous systematic literature review and updated the search to screen articles published from January 2017 to January 2020. We included randomized controlled trials and cluster randomized controlled trials that compared a CDS intervention with and without specific factors. We used random effects meta-regression procedures to analyze clinician behavior for the aggregate effects. The theoretical model was the Unified Theory of Acceptance and Use of Technology (UTAUT) model with motivational control. RESULTS: Thirty-four studies were included. The meta-regression models identified the importance of effort expectancy (estimated coefficient = -0.162; P = .0003); facilitating conditions (estimated coefficient = 0.094; P = .013); and performance expectancy with motivational control (estimated coefficient = 1.029; P = .022). Each of these factors created a significant impact on clinician behavior. The meta-regression model with the multivariate analysis explained a large amount of the heterogeneity across studies (R2 = 88.32%). DISCUSSION: Three positive factors were identified: low effort to use, low controllability, and providing more infrastructure and implementation strategies to support the CDS. The multivariate analysis suggests that passive CDS could be effective if users believe the CDS is useful and/or social expectations to use the CDS intervention exist. CONCLUSIONS: Overall, a modified UTAUT model that includes motivational control is an appropriate model to understand psychological factors associated with CDS effectiveness and to guide CDS design, implementation, and optimization. Siru Liu, Thomas J. Reese, Kensaku Kawamoto, Guilherme Del Fiol, Charlene R. Weir |
J. Am. Medical Informatics Assoc. | 1 |
| 2020 | Can the UTAUT Model Characterize Clinical Decision Support?
Siru Liu, Thomas J. Reese, Kensaku Kawamoto, Guilherme Del Fiol, Charlene R. Weir |
AMIA | 1 |
| 2019 | Clinical Informatics Course: Doctor of Medicine Student Perceptions
Siru Liu, Wenyao Annie Wu |
AMIA | 2 |
| 2019 | Using Natural Language Processing to improve EHR Structured Data-based Surgical Site Infection Surveillance
Jianlin Shi, Siru Liu, Liese C. Pruitt, Carolyn Luppens, Jeffrey P. Ferraro, Adi V. Gundlapalli, Wendy W. Chapman, Brian T. Bucher |
AMIA | 2 |
| 2018 | Detection of Healthcare-Associated Infections Using Electronic Health Record Data
Siru Liu, Jeffrey P. Ferraro, Adi V. Gundlapalli, Wendy W. Chapman, Brian T. Bucher |
AMIA | 1 |
| 2018 | Factors Affecting the Use of Electronic Health Records by Physician: A Pilot Study
Siru Liu |
AMIA | 1 |