EDBT 2026 Demo / reviewers in the wild / expert
Michael T. M. Baiocchi
dblp:228/8719
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-7571-5268ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing best practices in large language model research: an application to repeat promptingabstractOBJECTIVES: We aimed to demonstrate the importance of establishing best practices in large language model research, using repeat prompting as an illustrative example. MATERIALS AND METHODS: Using data from a prior study investigating potential model bias in peer review of medical abstracts, we compared methods that ignore correlation in model outputs from repeated prompting with a random effects method that accounts for this correlation. RESULTS: High correlation within groups was found when repeatedly prompting the model, with intraclass correlation coefficient of 0.69. Ignoring the inherent correlation in the data led to over 100-fold inflation of effective sample size. After appropriately accounting for this issue, the authors' results reverse from a small but highly significant finding to no evidence of model bias. DISCUSSION: The establishment of best practices for LLM research is urgently needed, as demonstrated in this case where accounting for repeat prompting in analyses was critical for accurate study conclusions. Robert Gallo, Michael T. M. Baiocchi, Thomas Savage, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | Monitoring strategies for continuous evaluation of deployed clinical prediction models
Grace Y. E. Kim, Conor K. Corbin, François Grolleau, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Biomed. Informatics | 4 |
| 2022 | How to Avoid Incorrect Clinical Machine Learning Model Performance Estimates When Class Labels Are Only Partially Observed
Conor K. Corbin, Michael T. M. Baiocchi, Jonathan H. Chen |
AMIA | 2 |
| 2021 | Randomized user testing of recommender system clinical decision support
Andre Kumar, Rachael C. Aikens, Jason Horn, Lisa Shieh, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
AMIA | 6 |
| 2021 | Integrating Evaluations of Predictive Algorithm-Driven Interventions into Clinical Workflows with the Dynamic Discontinuity Deployment Design
Ben J. Marafino, Alejandro Schuler, Vincent X. Liu, Art B. Owen, Gabriel J. Escobar, Michael T. M. Baiocchi |
AMIA | 6 |
| 2021 | Predicting Level of Care for Emergency Hospital Admissions to Optimize Triage
Nicolai P. Ostberg, Conor K. Corbin, Tiffany Eulalio, Gautam Machiraju, Ben J. Marafino, Michael T. M. Baiocchi, Christian Rose, Jonathan H. Chen |
AMIA | 7 |
| 2021 | Developing machine learning models to personalize care levels among emergency room patients for hospital admissionabstractOBJECTIVE: To develop prediction models for intensive care unit (ICU) vs non-ICU level-of-care need within 24 hours of inpatient admission for emergency department (ED) patients using electronic health record data. MATERIALS AND METHODS: Using records of 41 654 ED visits to a tertiary academic center from 2015 to 2019, we tested 4 algorithms-feed-forward neural networks, regularized regression, random forests, and gradient-boosted trees-to predict ICU vs non-ICU level-of-care within 24 hours and at the 24th hour following admission. Simple-feature models included patient demographics, Emergency Severity Index (ESI), and vital sign summary. Complex-feature models added all vital signs, lab results, and counts of diagnosis, imaging, procedures, medications, and lab orders. RESULTS: The best-performing model, a gradient-boosted tree using a full feature set, achieved an AUROC of 0.88 (95%CI: 0.87-0.89) and AUPRC of 0.65 (95%CI: 0.63-0.68) for predicting ICU care need within 24 hours of admission. The logistic regression model using ESI achieved an AUROC of 0.67 (95%CI: 0.65-0.70) and AUPRC of 0.37 (95%CI: 0.35-0.40). Using a discrimination threshold, such as 0.6, the positive predictive value, negative predictive value, sensitivity, and specificity were 85%, 89%, 30%, and 99%, respectively. Vital signs were the most important predictors. DISCUSSION AND CONCLUSIONS: Undertriaging admitted ED patients who subsequently require ICU care is common and associated with poorer outcomes. Machine learning models using readily available electronic health record data predict subsequent need for ICU admission with good discrimination, substantially better than the benchmarking ESI system. The results could be used in a multitiered clinical decision-support system to improve ED triage. Conor K. Corbin, Tiffany Eulalio, Nicolai P. Ostberg, Gautam Machiraju, Ben J. Marafino, Michael T. M. Baiocchi, Christian Rose, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Machine learning for initial insulin estimation in hospitalized patientsabstractOBJECTIVE: The study sought to determine whether machine learning can predict initial inpatient total daily dose (TDD) of insulin from electronic health records more accurately than existing guideline-based dosing recommendations. MATERIALS AND METHODS: Using electronic health records from a tertiary academic center between 2008 and 2020 of 16,848 inpatients receiving subcutaneous insulin who achieved target blood glucose control of 100-180 mg/dL on a calendar day, we trained an ensemble machine learning algorithm consisting of regularized regression, random forest, and gradient boosted tree models for 2-stage TDD prediction. We evaluated the ability to predict patients requiring more than 6 units TDD and their point-value TDDs to achieve target glucose control. RESULTS: The method achieves an area under the receiver-operating characteristic curve of 0.85 (95% confidence interval [CI], 0.84-0.87) and area under the precision-recall curve of 0.65 (95% CI, 0.64-0.67) for classifying patients who require more than 6 units TDD. For patients requiring more than 6 units TDD, the mean absolute percent error in dose prediction based on standard clinical calculators using patient weight is in the range of 136%-329%, while the regression model based on weight improves to 60% (95% CI, 57%-63%), and the full ensemble model further improves to 51% (95% CI, 48%-54%). DISCUSSION: Owingto the narrow therapeutic window and wide individual variability, insulin dosing requires adaptive and predictive approaches that can be supported through data-driven analytic tools. CONCLUSIONS: Machine learning approaches based on readily available electronic medical records can discriminate which inpatients will require more than 6 units TDD and estimate individual doses more accurately than standard guidelines and practices. Ivana Jankovic, Laurynas Kalesinskas, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | OrderRex clinical user testing: a randomized trial of recommender system decision support on simulated casesabstractOBJECTIVE: To assess usability and usefulness of a machine learning-based order recommender system applied to simulated clinical cases. MATERIALS AND METHODS: 43 physicians entered orders for 5 simulated clinical cases using a clinical order entry interface with or without access to a previously developed automated order recommender system. Cases were randomly allocated to the recommender system in a 3:2 ratio. A panel of clinicians scored whether the orders placed were clinically appropriate. Our primary outcome included the difference in clinical appropriateness scores. Secondary outcomes included total number of orders, case time, and survey responses. RESULTS: Clinical appropriateness scores per order were comparable for cases randomized to the order recommender system (mean difference -0.11 order per score, 95% CI: [-0.41, 0.20]). Physicians using the recommender placed more orders (median 16 vs 15 orders, incidence rate ratio 1.09, 95%CI: [1.01-1.17]). Case times were comparable with the recommender system. Order suggestions generated from the recommender system were more likely to match physician needs than standard manual search options. Physicians used recommender suggestions in 98% of available cases. Approximately 95% of participants agreed the system would be useful for their workflows. DISCUSSION: User testing with a simulated electronic medical record interface can assess the value of machine learning and clinical decision support tools for clinician usability and acceptance before live deployments. CONCLUSIONS: Clinicians can use and accept machine learned clinical order recommendations integrated into an electronic order entry interface in a simulated setting. The clinical appropriateness of orders entered was comparable even when supported by automated recommendations. Andre Kumar, Rachael C. Aikens, Jason Hom, Lisa Shieh, Jonathan Chiang, David Morales, Divya Saini, Mark A. Musen, Michael T. M. Baiocchi, Russ B. Altman, Mary K. Goldstein, Steven M. Asch, Jonathan H. Chen |
J. Am. Medical Informatics Assoc. | 9 |
| 2018 | An evaluation of clinical order patterns machine-learned from clinician cohorts stratified by patient mortality outcomesabstractEvaluate the quality of clinical order practice patterns machine-learned from clinician cohorts stratified by patient mortality outcomes. Inpatient electronic health records from 2010 to 2013 were extracted from a tertiary academic hospital. Clinicians (n = 1822) were stratified into low-mortality (21.8%, n = 397) and high-mortality (6.0%, n = 110) extremes using a two-sided P-value score quantifying deviation of observed vs. expected 30-day patient mortality rates. Three patient cohorts were assembled: patients seen by low-mortality clinicians, high-mortality clinicians, and an unfiltered crowd of all clinicians (n = 1046, 1046, and 5230 post-propensity score matching, respectively). Predicted order lists were automatically generated from recommender system algorithms trained on each patient cohort and evaluated against (i) real-world practice patterns reflected in patient cases with better-than-expected mortality outcomes and (ii) reference standards derived from clinical practice guidelines. Across six common admission diagnoses, order lists learned from the crowd demonstrated the greatest alignment with guideline references (AUROC range = 0.86–0.91), performing on par or better than those learned from low-mortality clinicians (0.79–0.84, P < 10−5) or manually-authored hospital order sets (0.65–0.77, P < 10−3). The same trend was observed in evaluating model predictions against better-than-expected patient cases, with the crowd model (AUROC mean = 0.91) outperforming the low-mortality model (0.87, P < 10−16) and order set benchmarks (0.78, P < 10−35). Whether machine-learning models are trained on all clinicians or a subset of experts illustrates a bias-variance tradeoff in data usage. Defining robust metrics to assess quality based on internal (e.g. practice patterns from better-than-expected patient cases) or external reference standards (e.g. clinical practice guidelines) is critical to assess decision support content. Learning relevant decision support content from all clinicians is as, if not more, robust than learning from a select subgroup of clinicians favored by patient outcomes. Jason K. Wang, Jason Hom, Santhosh Balasubramanian, Alejandro Schuler, Nigam H. Shah, Mary K. Goldstein, Michael T. M. Baiocchi, Jonathan H. Chen |
J. Biomed. Informatics | 7 |