Naveen Muthu

dblp:200/4594 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-8259-6965ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 10 since 2021
YearPublicationVenuePosition
2025 Human performance evaluation of a pediatric artificial intelligence sepsis model
abstract
OBJECTIVE: To assess the influence of an implemented artificial intelligence model predicting pediatric sepsis (defined by IPSO-Improving Pediatric Sepsis Outcomes collaborative) in the emergency department (ED) on human performance measures. MATERIALS AND METHODS: Two ED sites within a large pediatric health system in the Southeastern United States between January 1, 2021 and April 1, 2024. We interviewed ED providers and nurses within 72 hours of caring for a patient identified as potentially having sepsis by the predictive model. Thematic analysis of qualitative data was combined with electronic health record queries to assess measures of human performance, including situation awareness, explainability, human-computer agreement, workload, trust, automation bias, and relationship between staff and patients. RESULTS: We interviewed 40 clinicians. Participants found that the sepsis alert improved situation awareness, leading to changes in patient care management, resource allocation, and/or monitoring. Participants reported an average trust in the model-based alert of 3.8/5. Only 28% (555/1977) of sepsis huddles were done without alert firing, suggesting some automation bias. Treatment with antibiotics for IPSO sepsis cases was similar pre- and post-intervention without a huddle (9.3% vs 10.5%), though treatment doubled with huddle intervention (22.7%). NASA Task Load Index increased from 43 to 57 post-intervention. There was no report of adverse relationships with patients post-intervention. DISCUSSION: Human performance appeared to be generally positive with improved situation awareness and satisfaction with the alert-driven huddle. However, there was some evidence of automation bias and a slight increase in workload with the intervention. CONCLUSION: This study demonstrates the feasibility of evaluating multiple dimensions of human performance using a mixed methods approach for an AI model implemented in clinical practice. Future studies should aim to reduce the measurement burden of human performance metrics associated with AI implementation in acute care settings and assess the correlation between human performance measures and clinical outcomes.
Swaminathan Kandaswamy, Naveen Muthu, Nikolay Braykov, Rebekah Carter, Reena Blanco, Thuy Bui, Evan Orenstein, Mark V. Mai
J. Am. Medical Informatics Assoc.2
2025 Early clinical evaluation of a vendor developed pediatric artificial intelligence sepsis model in the emergency department
abstract
OBJECTIVE: To conduct an independent external validation of an implemented vendor-developed emergency department (ED) pediatric sepsis predictive model. MATERIALS AND METHODS: We performed a retrospective cross-sectional study within 2 ED sites of a large pediatric health system between January 1, 2021 and April 1, 2024. A nurse-facing interruptive alert appeared when the model score exceeded the threshold, triggering clinicians to call a sepsis huddle. We compared model predictive performance with vendor-reported performance using definitions that accounted for model threshold and alert timing in clinical practice. Care processes and patient outcome measures included time to first antibiotics, time to first fluid bolus, 30-day mortality, ED to ICU admission rate, and ICU free days. RESULTS: The pre-intervention cohort consisted of 268 102 ED visits with 741 (0.28%) sepsis cases. The post-intervention cohort consisted of 331 061 ED visits with 1114 (0.34%) sepsis cases. Model predictive performance dropped from vendor-reported performance. Mean time to first antibiotic decreased from 112 to 102 minutes (P = .05, 95% confidence interval of difference, -19.1 to 0.1) and time to first bolus decreased by 16.7 minutes (P = .03, 95% confidence interval difference, -31.8 to -1.5) after the intervention. Decreases in 30-day mortality (6% [45/741] to 4% [52/1114]); ED to ICU admissions (87% [646/741] to 84% [941/1114]), and ICU free days (6 to 5) after the intervention did not meet statistical significance. DISCUSSION: Implementing the model led to significant reductions in time to fluid bolus and borderline decreases in time to antibiotics, with non-significant changes in mortality and ICU metrics. When implementing an externally developed model, local workflows, documentation patterns, and patient populations make it challenging to generalize published or reported model performance metrics to real world performance. CONCLUSION: When tailoring a vendor-developed pediatric ED sepsis model for real-world usage, predictive performance differed substantially. Post-implementation we found improvements in care process measures, suggesting such models may benefit sepsis care when adapted for specific clinical workflows.
Swaminathan Kandaswamy, Evan Orenstein, Naveen Muthu, Andrea McCarter, Nikolay Braykov, Jonathan M. Beus, Edwin Ray, Tal Senior, Sara Brown, Rebekah Carter, Marybeth Gleeson, Hannah Thummel, John Cheng, Thuy Bui, Reena Blanco, Kiran Hebbar, James Fortenberry, Srikant Iyer, Mark V. Mai
J. Am. Medical Informatics Assoc.3
2025 Alert design in the real world: a cross-sectional analysis of interruptive alerting at 9 academic pediatric health systems
abstract
OBJECTIVE: To assess the prevalence of recommended design elements in implemented electronic health record (EHR) interruptive alerts across pediatric care settings. MATERIALS AND METHODS: We conducted a 3-phase mixed-methods cross-sectional study. Phase 1 involved developing a codebook for alert content classification. Phase 2 identified the most frequently interruptive alerts at participating sites. Phase 3 applied the codebook to classify alerts. Inter-rater reliability (IRR) for the codebook and descriptive statistics for alert design contents were reported. RESULTS: We classified alert content on design elements such as the rationale for the alert's appearance, the hazard of ignoring it, directive versus informational content, administrative purpose, and whether it aligned with one of the Institute of Medicine's (IOM) domains of healthcare quality. Most design elements achieved an IRR above 0.7, with the exceptions for identifying directive content outside of an alert (IRR 0.58) and whether an alert was for administrative purposes only (IRR 0.36). IRR was poor for all IOM domains except equity. Institutions varied widely in the number of unique alerts and their designs. 78% of alerts stated their purpose, over half were directive, and 13% were informational. Only 2%-20% of alerts explained the consequences of inaction. DISCUSSION: This study raises important questions about the optimal balance of alert functions and desirable features of alert representation. CONCLUSION: Our study provides the first multi-center analysis of EHR alert design elements in pediatric care settings, revealing substantial variation in content and design. These findings underline the need for future research to experimentally explore EHR alert design best practices to improve efficiency and effectiveness.
Swaminathan Kandaswamy, Julia K. W. Yarahuan, Elizabeth A. Dobler, Matthew J. Molloy, Lindsey A. Knake, Sean Hernandez, Anne A Fallon, Lauren M. Hess, Allison B. McCoy, Regine M. Fortunov, Eric S. Kirkendall, Naveen Muthu, Evan Orenstein, Adam C. Dziorny, Juan D. Chaparro
J. Am. Medical Informatics Assoc.12
2023 Clinical decision support with a comprehensive in-EHR patient tracking system improves genetic testing follow up
abstract
OBJECTIVE: We sought to develop and evaluate an electronic health record (EHR) genetic testing tracking system to address the barriers and limitations of existing spreadsheet-based workarounds. MATERIALS AND METHODS: We evaluated the spreadsheet-based system using mixed effects logistic regression to identify factors associated with delayed follow up. These factors informed the design of an EHR-integrated genetic testing tracking system. After deployment, we assessed the system in 2 ways. We analyzed EHR access logs and note data to assess patient outcomes and performed semistructured interviews with users to identify impact of the system on work. RESULTS: We found that patient-reported race was a significant predictor of documented genetic testing follow up, indicating a possible inequity in care. We implemented a CDS system including a patient data capture form and management dashboard to facilitate important care tasks. The system significantly sped review of results and significantly increased documentation of follow-up recommendations. Interviews with key system users identified a range of sociotechnical factors (ie, tools, tasks, collaboration) that contribute to safer and more efficient care. DISCUSSION: Our new tracking system ended decades of workarounds for identifying and communicating test results and improved clinical workflows. Interview participants related that the system decreased cognitive and time burden which allowed them to focus on direct patient interaction. CONCLUSION: By assembling a multidisciplinary team, we designed a novel patient tracking system that improves genetic testing follow up. Similar approaches may be effective in other clinical settings.
Ian M. Campbell, Dean Karavite, Morgan L. McManus, Fred C. Cusick, David C. Junod, Sarah E. Sheppard, Eli M. Lourie, Eric D. Shelov, Hakon Hakonarson, Anthony A. Luberti, Naveen Muthu, Robert Grundmeier
J. Am. Medical Informatics Assoc.11
2022 Quantifying Clinical Decision Support Across Pediatric Intensive Care Units
Alex Clark, Naveen Muthu, Mark V. Mai, Evan Orenstein, Swaminathan Kandaswamy, Kathleen Fear, Adam C. Dziorny
AMIA2
2021 Estimating Early Warning System Accuracy Prior to Implementation
Lusha Cao, Gerald P. Shaeffer, Meghan Galligan, Fuchiang R. Tsui, Robert Grundmeier, Vinay Nadkarni, Robert Sutton, Christopher P. Bonafide, Naveen Muthu
AMIA9
2021 Clinical Decision Support in the Pediatric ICU: A Multi-Institution Survey
Adam C. Dziorny, Julia A. Heneghan, Moodakare A. Bhat, Dean Karavite, L. Nelson Sanchez-Pinto, J. J. McArthur, Naveen Muthu
AMIA7
2021 Quantifying Changes in Resident-Patient Interactions During the COVID-19 Pandemic Using EHR Audit Logs
Mark V. Mai, Naveen Muthu, Bryn Carroll, Anna Costello, Dan West, Adam C. Dziorny
AMIA2
2021 Is your clinical decision support moving the needle on outcomes that matter? Novel software for evaluating quality improvement initiatives
Evan Orenstein, Naveen Muthu, Marc Tobias
AMIA2
2021 Alert burden in pediatric hospitals: a cross-sectional analysis of six academic pediatric health systems using novel metrics
abstract
BACKGROUND: Excessive electronic health record (EHR) alerts reduce the salience of actionable alerts. Little is known about the frequency of interruptive alerts across health systems and how the choice of metric affects which users appear to have the highest alert burden. OBJECTIVE: (1) Analyze alert burden by alert type, care setting, provider type, and individual provider across 6 pediatric health systems. (2) Compare alert burden using different metrics. MATERIALS AND METHODS: We analyzed interruptive alert firings logged in EHR databases at 6 pediatric health systems from 2016-2019 using 4 metrics: (1) alerts per patient encounter, (2) alerts per inpatient-day, (3) alerts per 100 orders, and (4) alerts per unique clinician days (calendar days with at least 1 EHR log in the system). We assessed intra- and interinstitutional variation and how alert burden rankings differed based on the chosen metric. RESULTS: Alert burden varied widely across institutions, ranging from 0.06 to 0.76 firings per encounter, 0.22 to 1.06 firings per inpatient-day, 0.98 to 17.42 per 100 orders, and 0.08 to 3.34 firings per clinician day logged in the EHR. Custom alerts accounted for the greatest burden at all 6 sites. The rank order of institutions by alert burden was similar regardless of which alert burden metric was chosen. Within institutions, the alert burden metric choice substantially affected which provider types and care settings appeared to experience the highest alert burden. CONCLUSION: Estimates of the clinical areas with highest alert burden varied substantially by institution and based on the metric used.
Evan Orenstein, Swaminathan Kandaswamy, Naveen Muthu, Juan D. Chaparro, Philip Hagedorn, Adam C. Dziorny, Adam Moses, Sean Hernandez, Amina Khan, Hannah B. Huth, Jonathan M. Beus, Eric S. Kirkendall
J. Am. Medical Informatics Assoc.3
2020 It's not just a last mile problem: Partnering with process improvement and human factors to integrate machine learning into healthcare delivery
Ron C. Li, Naveen Muthu, Margaret Smith, Swaminathan Kandaswamy, Jonathan H. Chen
AMIA2
2020 Clinical Decision Support for Health Maintenance Interventions in Acute Care Settings: Three Approaches to Promoting Influenza Vaccine
Evan Orenstein, Juan D. Chaparro, Emily C. Webber, Naveen Muthu
AMIA4
2020 A maximum likelihood approach to electronic health record phenotyping using positive and unlabeled patients
abstract
OBJECTIVE: Phenotyping patients using electronic health record (EHR) data conventionally requires labeled cases and controls. Assigning labels requires manual medical chart review and therefore is labor intensive. For some phenotypes, identifying gold-standard controls is prohibitive. We developed an accurate EHR phenotyping approach that does not require labeled controls. MATERIALS AND METHODS: Our framework relies on a random subset of cases, which can be specified using an anchor variable that has excellent positive predictive value and sensitivity independent of predictors. We proposed a maximum likelihood approach that efficiently leverages data from the specified cases and unlabeled patients to develop logistic regression phenotyping models, and compare model performance with existing algorithms. RESULTS: Our method outperformed the existing algorithms on predictive accuracy in Monte Carlo simulation studies, application to identify hypertension patients with hypokalemia requiring oral supplementation using a simulated anchor, and application to identify primary aldosteronism patients using real-world cases and anchor variables. Our method additionally generated consistent estimates of 2 important parameters, phenotype prevalence and the proportion of true cases that are labeled. DISCUSSION: Upon identification of an anchor variable that is scalable and transferable to different practices, our approach should facilitate development of scalable, transferable, and practice-specific phenotyping models. CONCLUSIONS: Our proposed approach enables accurate semiautomated EHR phenotyping with minimal manual labeling and therefore should greatly facilitate EHR clinical decision support and research.
Lingjiao Zhang, Xiruo Ding, Yanyuan Ma, Naveen Muthu, Imran Ajmal, Jason H. Moore, Daniel S. Herman
J. Am. Medical Informatics Assoc.4
2019 Variability in User Response to Custom Alerts in the Electronic Health Record: An Observational Study
Naveen Muthu, Eric D. Shelov, Marc Tobias, Dean Karavite, Evan Orenstein, Robert Grundmeier
AMIA1
2019 Reduction in Severe Ordering Errors of Blood Products in Pediatrics through Formative and Summative Usability Testing
Evan Orenstein, Jeanne Boudreaux, Margo Rollins, Christy Bryant, Dean Karavite, Naveen Muthu, Jessica Hike, Herb Williams, Alexis B. Carter, Cassandra Josephson
AMIA7
2018 Defining Clinically Meaningful Documentation Discrepancies during Transfer from the Pediatric Intensive Care Unit to Medical Services
Daria Ferro, Naveen Muthu, Christopher P. Bonafide, Evan Orenstein
AMIA2
2018 Preemptive Clinical Decision Support: Delivering Precise Information Clinicians Will Actually Use to Prevent Harm
James M. Hoffman, Josh F. Peterson, Naveen Muthu, Henry M. Dunnenberger
AMIA3
2018 Surveillance Methods for EHR Safety Hazards: A Panel Discussion
Naveen Muthu, Allan Fong, Daria Ferro, Evan Orenstein
AMIA1
2018 Functional Analysis of Written Communication Needs for Inpatient Providers
Evan Orenstein, Naveen Muthu, Subha L. Airan-Javia
AMIA2
2018 Decision Support for Decision Support: A Novel System to Prioritize Improvement Efforts, Identify Safety Hazards, and Measure Improvement
Marc Tobias, Evan Orenstein, Naveen Muthu
AMIA3
2016 Bridging the Gap Between Public Health and Clinical Provider: The PHRASE Interoperable Platform for Improving Decision Support
Marc Tobias, Naveen Muthu, Robert Grundmeier
AMIA2