VLDB 2026 Research / reviewers in the wild / expert
Katherine E. Brown
dblp:239/1024
· DBLP profile ↗
7ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0003-4443-8541ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 6 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gaps in artificial intelligence research for rural health in the United States: a scoping reviewabstractOBJECTIVE: Artificial intelligence (AI) has impacted healthcare at urban and academic medical centers in the US. There are concerns, however, that the promise of AI may not be realized in rural communities. This scoping review aims to determine the extent of AI research in the rural US. MATERIALS AND METHODS: We conducted a scoping review following the PRISMA guidelines. We included peer-reviewed, original research studies indexed in PubMed, Embase, and WebOfScience after January 1, 2010 and through April 29, 2025. Studies were required to discuss the development, implementation, or evaluation of AI tools in rural US healthcare, including frameworks that help facilitate AI development (eg, data warehouses). RESULTS: Our search strategy found 26 studies meeting inclusion criteria after full text screening with 14 papers discussing predictive AI models and 12 papers discussing data or research infrastructure. AI models most often targeted resource allocation and distribution. Few studies explored model deployment and impact. Half noted the lack of data and analytic resources as a limitation. None of the studies discussed examples of generative AI being trained, evaluated, or deployed in a rural setting. DISCUSSION: Practical limitations may be influencing and limiting the types of AI models evaluated in the rural US. Validation of tools in the rural US was underwhelming. CONCLUSION: With few studies moving beyond AI model design and development stages, there are clear gaps in our understanding of how to reliably validate, deploy, and sustain AI models in rural settings to advance health in all communities. Katherine E. Brown, Sharon E. Davis |
J. Am. Medical Informatics Assoc. | 1 |
| 2026 | Auditor models to suppress poor artificial intelligence predictions can improve human-artificial intelligence collaborative performanceabstractOBJECTIVE: Healthcare decisions are increasingly made with the assistance of machine learning (ML). ML has been known to have unfairness-inconsistent outcomes across subpopulations. Clinicians interacting with these systems can perpetuate such unfairness by overreliance. Recent work exploring ML suppression-silencing predictions based on auditing the ML-shows promise in mitigating performance issues originating from overreliance. This study aims to evaluate the impact of suppression on collaboration fairness and evaluate ML uncertainty as desiderata to audit the ML. MATERIALS AND METHODS: We used data from the Vanderbilt University Medical Center electronic health record (n = 58 817) and the MIMIC-IV-ED dataset (n = 363 145) to predict likelihood of death or intensive care unit transfer and likelihood of 30-day readmission using gradient-boosted trees and an artificially high-performing oracle model. We derived clinician decisions directly from the dataset and simulated clinician acceptance of ML predictions based on previous empirical work on acceptance of clinical decision support alerts. We measured performance as area under the receiver operating characteristic curve and algorithmic fairness using absolute averaged odds difference. RESULTS: When the ML outperforms humans, suppression outperforms the human alone (P < 8.2 × 10-6) and at least does not degrade fairness. When the human outperforms the ML, the human is either fairer than suppression (P < 8.2 × 10-4) or there is no statistically significant difference in fairness. Incorporating uncertainty quantification into suppression approaches can improve performance. CONCLUSION: Suppression of poor-quality ML predictions through an auditor model shows promise in improving collaborative human-AI performance and fairness. Katherine E. Brown, Jesse O. Wrenn, Nicholas J. Jackson, Michael R. Cauley, Benjamin X. Collins, Laurie L. Novak, Bradley A. Malin, Jessica S. Ancker |
J. Am. Medical Informatics Assoc. | 1 |
| 2026 | Factors influencing the effectiveness of artificial intelligence-assisted decision-making in medicine: a scoping reviewabstractOBJECTIVES: Research on artificial intelligence (AI)-based clinical decision-support (AI-CDS) systems has returned mixed results. Sometimes providing AI-CDS to a clinician will improve decision-making performance, sometimes it will not, and it is not always clear why. This scoping review seeks to clarify existing evidence by identifying clinician-level and technology design factors that impact the effectiveness of AI-assisted decision-making in medicine. MATERIALS AND METHODS: We searched MEDLINE, Web of Science, and Embase for peer-reviewed papers that studied factors impacting the effectiveness of AI-CDS. We identified the factors studied and their impact on 3 outcomes: clinicians' attitudes toward AI, their decisions (eg, acceptance rate of AI recommendations), and their performance when utilizing AI-CDS. RESULTS: We retrieved 5850 articles and included 45. Four clinician-level and technology design factors were commonly studied. Expert clinicians may benefit less from AI-CDS than nonexperts, with some mixed results. Explainable AI increased clinicians' trust, but could also increase trust in incorrect AI recommendations, potentially harming human-AI collaborative performance. Clinicians' baseline attitudes toward AI predict their acceptance rates of AI recommendations. Of the 3 outcomes of interest, human-AI collaborative performance was most commonly assessed. DISCUSSION AND CONCLUSION: Few factors have been studied for their impact on the effectiveness of AI-CDS. Due to conflicting outcomes between studies, we recommend future work should leverage the concept of "appropriate trust" to facilitate more robust research on AI-CDS, aiming not to increase overall trust in or acceptance of AI but to ensure that clinicians accept AI recommendations only when trust in AI is warranted. Nicholas J. Jackson, Katherine E. Brown, Rachael Miller, Matthew Murrow, Michael R. Cauley, Benjamin X. Collins, Laurie L. Novak, Natalie C. Benda, Jessica S. Ancker |
J. Am. Medical Informatics Assoc. | 2 |
| 2026 | Community medical centers struggle to produce well-calibrated clinical prediction models: Data augmentation can helpabstractOBJECTIVE: Machine learning models (ML) often require localization to perform optimally in local populations. We hypothesize that smaller community healthcare centers may not have the necessary patient volume to facilitate localization based on statistical guidelines. This work investigates the ability for community medical centers to localize ML and performs a simulation study to evaluate synthetic data generation (SDG) to augment local data for recalibration. METHODS: We conducted an experiment using data from a real network of hospitals (two rural, one urban academic medical center) to predict 30-day unplanned hospital readmission and using data from a multi-site ICU dataset to simulate using synthetic data generation (SDG) in a network of hospitals of various sizes. We also performed a simulation study using data from a multi-site ICU dataset to evaluate the utility of SDG to augment local data volumes. RESULTS: In the real-world evaluation, the urban medical center met the guidelines for the number of samples for recalibration (Required: 14,224, Available: 42,303) and had the best calibrated model using local data (α=0.1,β=1.05; best: α=0,β=1). For the smaller sites, neither site had the samples required for recalibration (Site 1: Required: 16461, Available: 3187; Site 2: Required: 15299, Available: 905). In the simulation study, deep learning-based SDG was most effective at improving calibration performance. CONCLUSIONS: Connections to large medical centers are not enough to promote accurate ML at all sites within a healthcare system. Data augmentation and SDG may provide the necessary data volumes to enable local recalibration at smaller facilities. Katherine E. Brown, Bradley A. Malin, Sharon E. Davis |
J. Biomed. Informatics | 1 |
| 2025 | Large language models are less effective at clinical prediction tasks than locally trained machine learning modelsabstractOBJECTIVES: To determine the extent to which current large language models (LLMs) can serve as substitutes for traditional machine learning (ML) as clinical predictors using data from electronic health records (EHRs), we investigated various factors that can impact their adoption, including overall performance, calibration, fairness, and resilience to privacy protections that reduce data fidelity. MATERIALS AND METHODS: We evaluated GPT-3.5, GPT-4, and traditional ML (as gradient-boosting trees) on clinical prediction tasks in EHR data from Vanderbilt University Medical Center (VUMC) and MIMIC IV. We measured predictive performance with area under the receiver operating characteristic (AUROC) and model calibration using Brier Score. To evaluate the impact of data privacy protections, we assessed AUROC when demographic variables are generalized. We evaluated algorithmic fairness using equalized odds and statistical parity across race, sex, and age of patients. We also considered the impact of using in-context learning by incorporating labeled examples within the prompt. RESULTS: Traditional ML [AUROC: 0.847, 0.894 (VUMC, MIMIC)] substantially outperformed GPT-3.5 (AUROC: 0.537, 0.517) and GPT-4 (AUROC: 0.629, 0.602) (with and without in-context learning) in predictive performance and output probability calibration [Brier Score (ML vs GPT-3.5 vs GPT-4): 0.134 vs 0.384 vs 0.251, 0.042 vs 0.06 vs 0.219)]. DISCUSSION: Traditional ML is more robust than GPT-3.5 and GPT-4 in generalizing demographic information to protect privacy. GPT-4 is the fairest model according to our selected metrics but at the cost of poor model performance. CONCLUSION: These findings suggest that non-fine-tuned LLMs are less effective and robust than locally trained ML for clinical prediction tasks, but they are improving across releases. Katherine E. Brown, Chao Yan 0004, Xinmeng Zhang, Benjamin X. Collins, You Chen 0001, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Estimating Uncertainty in Deep Image Classification
Katherine E. Brown, Douglas A. Talbert |
AMIA | 1 |
| 2018 | Multi-Task Correlation-Based Feature Selection for Gene Expression Data
Katherine E. Brown, Douglas A. Talbert |
AMIA | 1 |