EDBT 2026 Demo / reviewers in the wild / expert
Armando Bedoya
dblp:214/4076
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-6496-7024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comparing ambient scribes: a randomized crossover clinical trial addressing ambient scribe technologies' impact on physician burnoutabstractOBJECTIVE: This study aims to compare the effectiveness of 2 ambient AI scribe technologies in reducing physician burnout, improving workflow satisfaction, and enhancing documentation efficiency through a randomized crossover trial. MATERIALS AND METHODS: An open-label randomized crossover trial involving 160 outpatient clinicians was conducted at a tertiary academic medical center. Volunteers were randomized to 2 groups of 80 with 2 crossover periods. We assessed workflow satisfaction (1-7 scale), burnout (Copenhagen Burnout Index), and efficiency metrics (eg, electronic health record time outside scheduled hours, documentation time, etc.). Data was analyzed using Wilcoxon signed-rank tests and generalized linear mixed models. RESULTS: Surveys from 136 respondents were analyzed. Clinicians reported greater improvements in satisfaction with product B (2.51 points on a 7-point scale) compared to product A (1.91 points; mean difference: 0.60, 95% CI: 0.32-0.90). Both tools reduced personal and work burnout scores, but differences between tools were not meaningful. Product B demonstrated greater reductions in average minutes-in-notes per day compared to product A (B - A = -3.19 minutes; 95% CI -4.87 to -1.50). No meaningful differences were observed in pajama time or patient-related burnout. DISCUSSION: Both tools improved workflow satisfaction and reduced burnout, with product B showing superior performance in satisfaction and documentation time. However, efficiency metrics like pajama time were largely unaffected, potentially due to participant selection bias and the study period's timing. CONCLUSION: Product B yielded greater satisfaction and time savings compared to product A, though both tools effectively reduced physician burnout and improved workflow satisfaction. Anand Chowdhury, Michele Casey, Jonathan Wilson, Kathryn I. Pollak, Benjamin Goldstein 0001, Armando Bedoya, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 6 |
| 2026 | A federated learning framework for ethical dynamic treatment allocation across heterogeneous hospitals
Xenia Konti, Nicoleta J. Economou-Zavlanos, Yi Shen 0011, Giorgos B. Stamou, Armando Bedoya, Michael J. Pencina, Chuan Hong, Michael M. Zavlanos |
J. Biomed. Informatics | 5 |
| 2025 | Application of unified health large language model evaluation framework to In-Basket message replies: bridging qualitative and quantitative assessmentsabstractOBJECTIVES: Large language models (LLMs) are increasingly utilized in healthcare, transforming medical practice through advanced language processing capabilities. However, the evaluation of LLMs predominantly relies on human qualitative assessment, which is time-consuming, resource-intensive, and may be subject to variability and bias. There is a pressing need for quantitative metrics to enable scalable, objective, and efficient evaluation. MATERIALS AND METHODS: We propose a unified evaluation framework that bridges qualitative and quantitative methods to assess LLM performance in healthcare settings. This framework maps evaluation aspects-such as linguistic quality, efficiency, content integrity, trustworthiness, and usefulness-to both qualitative assessments and quantitative metrics. We apply our approach to empirically evaluate the Epic In-Basket feature, which uses LLM to generate patient message replies. RESULTS: The empirical evaluation demonstrates that while Artificial Intelligence (AI)-generated replies exhibit high fluency, clarity, and minimal toxicity, they face challenges with coherence and completeness. Clinicians' manual decision to use AI-generated drafts correlates strongly with quantitative metrics, suggesting that quantitative metrics have the potential to reduce human effort in the evaluation process and make it more scalable. DISCUSSION: Our study highlights the potential of a unified evaluation framework that integrates qualitative and quantitative methods, enabling scalable and systematic assessments of LLMs in healthcare. Automated metrics streamline evaluation and monitoring processes, but their effective use depends on alignment with human judgment, particularly for aspects requiring contextual interpretation. As LLM applications expand, refining evaluation strategies and fostering interdisciplinary collaboration will be critical to maintaining high standards of accuracy, ethics, and regulatory compliance. CONCLUSION: Our unified evaluation framework bridges the gap between qualitative human assessments and automated quantitative metrics, enhancing the reliability and scalability of LLM evaluations in healthcare. While automated quantitative evaluations are not ready to fully replace qualitative human evaluations, they can be used to enhance the process and, with relevant benchmarks derived from the unified framework proposed here, they can be applied to LLM monitoring and evaluation of updated versions of the original technology evaluated using qualitative human standards. Chuan Hong, Anand Chowdhury, Anthony D. Sorrentino, Monica Agrawal, Armando Bedoya, Sophia Bessias, Nicoleta J. Economou-Zavlanos, Ian Wong, Christian Pean, Kathryn I. Pollak, Eric G. Poon, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Translating ethical and quality principles for the effective, safe and fair development, deployment and use of artificial intelligence technologies in healthcareabstractOBJECTIVE: The complexity and rapid pace of development of algorithmic technologies pose challenges for their regulation and oversight in healthcare settings. We sought to improve our institution's approach to evaluation and governance of algorithmic technologies used in clinical care and operations by creating an Implementation Guide that standardizes evaluation criteria so that local oversight is performed in an objective fashion. MATERIALS AND METHODS: Building on a framework that applies key ethical and quality principles (clinical value and safety, fairness and equity, usability and adoption, transparency and accountability, and regulatory compliance), we created concrete guidelines for evaluating algorithmic technologies at our institution. RESULTS: An Implementation Guide articulates evaluation criteria used during review of algorithmic technologies and details what evidence supports the implementation of ethical and quality principles for trustworthy health AI. Application of the processes described in the Implementation Guide can lead to algorithms that are safer as well as more effective, fair, and equitable upon implementation, as illustrated through 4 examples of technologies at different phases of the algorithmic lifecycle that underwent evaluation at our academic medical center. DISCUSSION: By providing clear descriptions/definitions of evaluation criteria and embedding them within standardized processes, we streamlined oversight processes and educated communities using and developing algorithmic technologies within our institution. CONCLUSIONS: We developed a scalable, adaptable framework for translating principles into evaluation criteria and specific requirements that support trustworthy implementation of algorithmic technologies in patient care and healthcare operations. Nicoleta J. Economou-Zavlanos, Sophia Bessias, Michael P. Cary, Armando Bedoya, Benjamin Goldstein 0001, John Eric Jelovsek, Cara O'Brien, Nancy Walden, Matthew Elmore, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Michael E. Lipkin, Afreen Idris Shariff, Michael Gao, David Leverenz, Ricardo Henao, David Y. Ming, David M. Gallagher, Michael J. Pencina, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | A framework for the oversight and local deployment of safe and high-quality prediction modelsabstractArtificial intelligence/machine learning models are being rapidly developed and used in clinical practice. However, many models are deployed without a clear understanding of clinical or operational impact and frequently lack monitoring plans that can detect potential safety signals. There is a lack of consensus in establishing governance to deploy, pilot, and monitor algorithms within operational healthcare delivery workflows. Here, we describe a governance framework that combines current regulatory best practices and lifecycle management of predictive models being used for clinical care. Since January 2021, we have successfully added models to our governance portfolio and are currently managing 52 models. Armando Bedoya, Nicoleta J. Economou-Zavlanos, Benjamin Goldstein 0001, Allison Young, John Eric Jelovsek, Cara O'Brien, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Erich Huang, Eric G. Poon, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 1 |
| 2021 | Looking for clinician involvement under the wrong lamp post: The need for collaboration measuresabstractIn a recent review published by Schwartz et al,1 an interdisciplinary team examines clinician involvement in machine learning research. The review includes 80 studies describing predictive clinical decision support systems (CDSSs) targeting clinicians for prognostic or treatment decision making in the hospital using electronic health record data. The objective of the review is to describe clinician involvement in these 80 studies and to map involvement across Stead’s 5 stages of system design.2 Unfortunately, the review makes 2 assumptions about interdisciplinary collaboration that undermine the analysis and interpretation of results. First, the review relies on a novel, highly constrained definition of clinician involvement. The constraints neglect substantial documentation of interdisciplinary collaboration, resulting in dramatic underestimates of collaboration. Second, the review misinterprets missing data. Studies that do not meet the constrained definition of clinician involvement are assumed to have been conducted without any clinician involvement. We highlight the weaknesses of... Mark P. Sendak, Michael Gao, William Ratliff, Marshall Nichols, Armando Bedoya, Cara O'Brien, Suresh Balu |
J. Am. Medical Informatics Assoc. | 5 |