EDBT 2026 Demo / reviewers in the wild / expert
Sharon E. Davis
dblp:186/5194
· DBLP profile ↗
25ranked-venue papers
14as first author
13since 2021 · last 2026
0000-0003-0792-8867ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 14 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Statistical methods to harmonize electronic health record data across healthcare systems: case study and lessons learnedabstractMOTIVATION: Although common data models for electronic health record (EHR) data can facilitate multi-site data organization and querying, the same medical event may still be coded differently between healthcare systems. In this paper, we present statistical methods to identify and mitigate coding discrepancies using summary-level data, and demonstrate these methods using data from two FDA Sentinel data partners: Kaiser Permanente Washington and Kaiser Permanente Northwest. RESULTS: We first characterize differences in coding patterns, then compute a code mapping matrix to harmonize data between systems. Our findings reveal significant heterogeneity in coded EHR data, even after adopting a common data model with the same coding system, highlighting the importance of data harmonization before downstream analyses. Our study also demonstrates the effectiveness of the data harmonization approaches, which provide a foundational data quality step to promote semantic interoperability, enhance data integration, and improve the integrity of study conclusions. AVAILABILITY AND IMPLEMENTATION: Computation prototypes, including R/Python codes and examples, are included in Section 7, available as supplementary data at Bioinformatics online and will be posted on GitHub upon publication. Yuqi Zhai, Xianshi Yu, Brian L. Hazlehurst, Denis B. Nyongesa, Daniel S. Sapp, Brian D. Williamson, David Carrell, Luesa Healy, Kara L. Cushing-Haugen, Jenna Wong, Shirley V. Wang, James S. Floyd, Kathleen Shattuck, Samuel McGown, Sarah Alam, José J. Hernández-Muñoz, Danijela Stojanovic, Sudha R. Raman, Sharon E. Davis, Tianxi Cai, Jennifer C. Nelson, Patrick J. Heagerty |
Bioinform. | 23 |
| 2026 | Gaps in artificial intelligence research for rural health in the United States: a scoping reviewabstractOBJECTIVE: Artificial intelligence (AI) has impacted healthcare at urban and academic medical centers in the US. There are concerns, however, that the promise of AI may not be realized in rural communities. This scoping review aims to determine the extent of AI research in the rural US. MATERIALS AND METHODS: We conducted a scoping review following the PRISMA guidelines. We included peer-reviewed, original research studies indexed in PubMed, Embase, and WebOfScience after January 1, 2010 and through April 29, 2025. Studies were required to discuss the development, implementation, or evaluation of AI tools in rural US healthcare, including frameworks that help facilitate AI development (eg, data warehouses). RESULTS: Our search strategy found 26 studies meeting inclusion criteria after full text screening with 14 papers discussing predictive AI models and 12 papers discussing data or research infrastructure. AI models most often targeted resource allocation and distribution. Few studies explored model deployment and impact. Half noted the lack of data and analytic resources as a limitation. None of the studies discussed examples of generative AI being trained, evaluated, or deployed in a rural setting. DISCUSSION: Practical limitations may be influencing and limiting the types of AI models evaluated in the rural US. Validation of tools in the rural US was underwhelming. CONCLUSION: With few studies moving beyond AI model design and development stages, there are clear gaps in our understanding of how to reliably validate, deploy, and sustain AI models in rural settings to advance health in all communities. Katherine E. Brown, Sharon E. Davis |
J. Am. Medical Informatics Assoc. | 2 |
| 2026 | Explainability in context: calibrating appropriate trust and reliance in artificial intelligenceabstractBACKGROUND AND SIGNIFICANCE: Predictive artificial intelligence (AI) promises to transform care delivery, enhance patient safety, and improve health outcomes. Realizing these benefits will require careful design, implementation, and monitoring strategies to avoid unintended consequences, including automation bias (i.e., erroneously favoring recommendations from automated systems). Automation bias is particularly concerning due to the variability of AI performance across time and populations, leading to predictions that may be variably incorrect, uncertain, or unfair. APPROACH: We advocate for an expanded view of explainable AI that uses contextual information to help end users calibrate appropriate levels of trust and reliance. We propose multiple levels of contextualization-model, setting, subpopulation, and patient-that together provide insight for clinicians to evaluate the reliability of individual predictions. This includes information about historical and in-the-moment AI performance, algorithmic fairness, and prediction uncertainty. CONCLUSION: We outline an approach to integrate context-based explanations into decision support workflows to aid clinician interpretation without adding cognitive burden. Sharon E. Davis, Megan E. Salwei |
J. Am. Medical Informatics Assoc. | 1 |
| 2026 | Community medical centers struggle to produce well-calibrated clinical prediction models: Data augmentation can helpabstractOBJECTIVE: Machine learning models (ML) often require localization to perform optimally in local populations. We hypothesize that smaller community healthcare centers may not have the necessary patient volume to facilitate localization based on statistical guidelines. This work investigates the ability for community medical centers to localize ML and performs a simulation study to evaluate synthetic data generation (SDG) to augment local data for recalibration. METHODS: We conducted an experiment using data from a real network of hospitals (two rural, one urban academic medical center) to predict 30-day unplanned hospital readmission and using data from a multi-site ICU dataset to simulate using synthetic data generation (SDG) in a network of hospitals of various sizes. We also performed a simulation study using data from a multi-site ICU dataset to evaluate the utility of SDG to augment local data volumes. RESULTS: In the real-world evaluation, the urban medical center met the guidelines for the number of samples for recalibration (Required: 14,224, Available: 42,303) and had the best calibrated model using local data (α=0.1,β=1.05; best: α=0,β=1). For the smaller sites, neither site had the samples required for recalibration (Site 1: Required: 16461, Available: 3187; Site 2: Required: 15299, Available: 905). In the simulation study, deep learning-based SDG was most effective at improving calibration performance. CONCLUSIONS: Connections to large medical centers are not enough to promote accurate ML at all sites within a healthcare system. Data augmentation and SDG may provide the necessary data volumes to enable local recalibration at smaller facilities. Katherine E. Brown, Bradley A. Malin, Sharon E. Davis |
J. Biomed. Informatics | 3 |
| 2025 | Emerging algorithmic bias: fairness drift as the next dimension of model maintenance and sustainabilityabstractOBJECTIVES: While performance drift of clinical prediction models is well-documented, the potential for algorithmic biases to emerge post-deployment has had limited characterization. A better understanding of how temporal model performance may shift across subpopulations is required to incorporate fairness drift into model maintenance strategies. MATERIALS AND METHODS: We explore fairness drift in a national population over 11 years, with and without model maintenance aimed at sustaining population-level performance. We trained random forest models predicting 30-day post-surgical readmission, mortality, and pneumonia using 2013 data from US Department of Veterans Affairs facilities. We evaluated performance quarterly from 2014 to 2023 by self-reported race and sex. We estimated discrimination, calibration, and accuracy, and operationalized fairness using metric parity measured as the gap between disadvantaged and advantaged groups. RESULTS: Our cohort included 1 739 666 surgical cases. We observed fairness drift in both the original and temporally updated models. Model updating had a larger impact on overall performance than fairness gaps. During periods of stable fairness, updating models at the population level increased, decreased, or did not impact fairness gaps. During periods of fairness drift, updating models restored fairness in some cases and exacerbated fairness gaps in others. DISCUSSION: This exploratory study highlights that algorithmic fairness cannot be assured through one-time assessments during model development. Temporal changes in fairness may take multiple forms and interact with model updating strategies in unanticipated ways. CONCLUSION: Equitable and sustainable clinical artificial intelligence deployments will require novel methods to monitor algorithmic fairness, detect emerging bias, and adopt model updates that promote fairness. Sharon E. Davis, Chad Dorn, Daniel J. Park, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | A machine learning framework to adjust for learning effects in medical device safety evaluationabstractOBJECTIVES: Traditional methods for medical device post-market surveillance often fail to accurately account for operator learning effects, leading to biased assessments of device safety. These methods struggle with non-linearity, complex learning curves, and time-varying covariates, such as physician experience. To address these limitations, we sought to develop a machine learning (ML) framework to detect and adjust for operator learning effects. MATERIALS AND METHODS: A gradient-boosted decision tree ML method was used to analyze synthetic datasets that replicate the complexity of clinical scenarios involving high-risk medical devices. We designed this process to detect learning effects using a risk-adjusted cumulative sum method, quantify the excess adverse event rate attributable to operator inexperience, and adjust for these alongside patient factors in evaluating device safety signals. To maintain integrity, we employed blinding between data generation and analysis teams. Synthetic data used underlying distributions and patient feature correlations based on clinical data from the Department of Veterans Affairs between 2005 and 2012. We generated 2494 synthetic datasets with widely varying characteristics including number of patient features, operators and institutions, and the operator learning form. Each dataset contained a hypothetical study device, Device B, and a reference device, Device A. We evaluated accuracy in identifying learning effects and identifying and estimating the strength of the device safety signal. Our approach also evaluated different clinically relevant thresholds for safety signal detection. RESULTS: Our framework accurately identified the presence or absence of learning effects in 93.6% of datasets and correctly determined device safety signals in 93.4% of cases. The estimated device odds ratios' 95% confidence intervals were accurately aligned with the specified ratios in 94.7% of datasets. In contrast, a comparative model excluding operator learning effects significantly underperformed in detecting device signals and in accuracy. Notably, our framework achieved 100% specificity for clinically relevant safety signal thresholds, although sensitivity varied with the threshold applied. DISCUSSION: A machine learning framework, tailored for the complexities of post-market device evaluation, may provide superior performance compared to standard parametric techniques when operator learning is present. CONCLUSION: Demonstrating the capacity of ML to overcome complex evaluative challenges, our framework addresses the limitations of traditional statistical methods in current post-market surveillance processes. By offering a reliable means to detect and adjust for learning effects, it may significantly improve medical device safety evaluation. Jejo Koola, Karthik Ramesh, Jialin Mao, Minyoung Ahn, Sharon E. Davis, Usha Govindarajulu, Amy Perkins, Dax M. Westerman, Henry Ssemaganda, Theodore Speroff, Lucila Ohno-Machado, Craig Ramsay, Art Sedrakyan, Frederic S. Resnic, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Detecting Opioid Use Disorder in Health Claims Data With Positive Unlabeled LearningabstractAccurate detection and prevalence estimation of behavioral health conditions, such as opioid use disorder (OUD), are crucial for identifying at-risk individuals, determining treatment needs, monitoring prevention and intervention efforts, and recruiting treatment-naive participants for clinical trials. The availability of extensive health data, combined with advancements in machine learning (ML) frameworks, has enabled researchers to employ various ML techniques to predict or identify OUD within patient health data. Ideally, we could directly estimate the prevalence, or the proportion of a population with a condition over time. However, underdiagnosis and undercoding of conditions in patient health records make it challenging to determine the true prevalence of these conditions and to identify at-risk patients with less severe conditions who are more likely to be missed. Consequently, patients without diagnoses may comprise positive and negative examples for a given condition. Treating all undiagnosed (uncoded) patients as negative when applying ML methods can introduce bias into models, affecting their predictive power. To address this issue, we employed Positive Unlabeled Learning Selected Not At Random (PULSNAR), a Positive and Unlabeled (PU) learning technique, to estimate the probability of a given patient having OUD during a time window and the overall population prevalence of OUD. In a sample of 3,342,044 commercially insured US patients with at least one opioid prescription filled, PULSNAR estimated that 5.08% of patients have a cumulative prevalence of OUD over a 2-5 a observation period, compared to the 1.35% with a recorded OUD diagnosis, with 73.5% of cases not diagnosed/coded. The prevalence estimates provided by PULSNAR are consistent with those reported in other studies. Fariha Moomtaheen, Scott A. Malec, Jeremy J. Yang, Cristian Bologa, Kristan Alexander Schneider, Yiliang Zhu 0001, Mauricio Tohen, Gerardo Villarreal, Douglas J. Perkins, Elliot M. Fielstein, Sharon E. Davis, Michael E. Matheny, Christophe G. Lambert |
IEEE J. Biomed. Health Informatics | 12 |
| 2024 | Sustainable deployment of clinical prediction tools - a 360° approach to model maintenanceabstractBACKGROUND: As the enthusiasm for integrating artificial intelligence (AI) into clinical care grows, so has our understanding of the challenges associated with deploying impactful and sustainable clinical AI models. Complex dataset shifts resulting from evolving clinical environments strain the longevity of AI models as predictive accuracy and associated utility deteriorate over time. OBJECTIVE: Responsible practice thus necessitates the lifecycle of AI models be extended to include ongoing monitoring and maintenance strategies within health system algorithmovigilance programs. We describe a framework encompassing a 360° continuum of preventive, preemptive, responsive, and reactive approaches to address model monitoring and maintenance from critically different angles. DISCUSSION: We describe the complementary advantages and limitations of these four approaches and highlight the importance of such a coordinated strategy to help ensure the promise of clinical AI is not short-lived. Sharon E. Davis, Peter J. Embí, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | A framework for understanding label leakage in machine learning for health careabstractINTRODUCTION: The pitfalls of label leakage, contamination of model input features with outcome information, are well established. Unfortunately, avoiding label leakage in clinical prediction models requires more nuance than the common advice of applying "no time machine rule." FRAMEWORK: We provide a framework for contemplating whether and when model features pose leakage concerns by considering the cadence, perspective, and applicability of predictions. To ground these concepts, we use real-world clinical models to highlight examples of appropriate and inappropriate label leakage in practice. RECOMMENDATIONS: Finally, we provide recommendations to support clinical and technical stakeholders as they evaluate the leakage tradeoffs associated with model design, development, and implementation decisions. By providing common language and dimensions to consider when designing models, we hope the clinical prediction community will be better prepared to develop statistically valid and clinically useful machine learning models. Sharon E. Davis, Michael E. Matheny, Suresh Balu, Mark P. Sendak |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | A Framework for Generating Synthetic Clinical Datasets with Learning Effects to Support Methods Development and Validation
Sharon E. Davis, Henry Ssemaganda, Jejo Koola, Jialin Mao, Dax M. Westerman, Theodore Speroff, Usha Govindarajulu, Craig Ramsay, Lucila Ohno-Machado, Frederic S. Resnic, Michael E. Matheny |
AMIA | 1 |
| 2022 | A Framework for Detecting Medical Device Safety Signals Confounded by Learning Effects Using Machine Learning
Jejo Koola, Jialin Mao, Sharon E. Davis, Henry Ssemaganda, Dax M. Westerman, Lucila Ohno-Machado, Frederic S. Resnic, Michael E. Matheny |
AMIA | 3 |
| 2022 | Disentangling and Characterizing Device Safety Signals and Learning Effects
Henry Ssemaganda, Frederic S. Resnic, Sharon E. Davis, Usha Govindarajulu, Jejo Koola, Jialin Mao, Dax M. Westerman, Theodore Speroff, Craig Ramsay, Art Sedrakyan, Lucila Ohno-Machado, Michael E. Matheny |
AMIA | 3 |
| 2021 | Disparities in Coded and Imputed Post-Traumatic Stress Disorder and Self-Harm Among US Veterans
Sharon E. Davis, Nicolas R. Lauve, Sharidan K. Parr, Daniel Park, Michael E. Matheny, Gerardo Villarreal, George Uhl, Yiliang Zhu 0001, Mauricio Tohen, Douglas J. Perkins, Christophe G. Lambert |
AMIA | 1 |
| 2020 | A Surveillance Framework for Monitoring and Updating Clinical Prediction Models
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 1 |
| 2020 | Detection of calibration drift in clinical prediction models to inform model updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Biomed. Informatics | 1 |
| 2019 | Comparison of Prediction Model Performance Updating Protocols: Using a Data-Driven Testing Procedure to Guide Updating
Sharon E. Davis, Robert A. Greevy Jr., Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
AMIA | 1 |
| 2019 | A nonparametric updating method to correct clinical prediction model driftabstractOBJECTIVE: Clinical prediction models require updating as performance deteriorates over time. We developed a testing procedure to select updating methods that minimizes overfitting, incorporates uncertainty associated with updating sample sizes, and is applicable to both parametric and nonparametric models. MATERIALS AND METHODS: We describe a procedure to select an updating method for dichotomous outcome models by balancing simplicity against accuracy. We illustrate the test's properties on simulated scenarios of population shift and 2 models based on Department of Veterans Affairs inpatient admissions. RESULTS: In simulations, the test generally recommended no update under no population shift, no update or modest recalibration under case mix shifts, intercept correction under changing outcome rates, and refitting under shifted predictor-outcome associations. The recommended updates provided superior or similar calibration to that achieved with more complex updating. In the case study, however, small update sets lead the test to recommend simpler updates than may have been ideal based on subsequent performance. DISCUSSION: Our test's recommendations highlighted the benefits of simple updating as opposed to systematic refitting in response to performance drift. The complexity of recommended updating methods reflected sample size and magnitude of performance drift, as anticipated. The case study highlights the conservative nature of our test. CONCLUSIONS: This new test supports data-driven updating of models developed with both biostatistical and machine learning approaches, promoting the transportability and maintenance of a wide array of clinical prediction models and, in turn, a variety of applications relying on modern prediction tools. Sharon E. Davis, Robert A. Greevy Jr., Christopher Fonnesbeck, Thomas A. Lasko, Colin G. Walsh, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Developing a Testing Procedure to Select Model Updating Methods
Sharon E. Davis, Robert A. Greevy Jr., Michael E. Matheny |
AMIA | 1 |
| 2018 | Development of an automated phenotyping algorithm for hepatorenal syndrome
Jejo Koola, Sharon E. Davis, Omar Al-Nimri, Sharidan K. Parr, Daniel Fabbri, Bradley A. Malin, Samuel B. Ho, Michael E. Matheny |
J. Biomed. Informatics | 2 |
| 2017 | Calibration Drift Among Regression and Machine Learning Models for Hospital Mortality
Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Michael E. Matheny |
AMIA | 1 |
| 2017 | Machine Learning Models to Predict Readmission for Patients with Cirrhosis
Jejo Koola, Aize Cao, Guanhua Chen 0002, Amy Perkins, Samuel B. Ho, Sharon E. Davis, Michael E. Matheny |
AMIA | 6 |
| 2017 | Calibration drift in regression and machine learning models for acute kidney injuryabstractOBJECTIVE: Predictive analytics create opportunities to incorporate personalized risk estimates into clinical decision support. Models must be well calibrated to support decision-making, yet calibration deteriorates over time. This study explored the influence of modeling methods on performance drift and connected observed drift with data shifts in the patient population. MATERIALS AND METHODS: Using 2003 admissions to Department of Veterans Affairs hospitals nationwide, we developed 7 parallel models for hospital-acquired acute kidney injury using common regression and machine learning methods, validating each over 9 subsequent years. RESULTS: Discrimination was maintained for all models. Calibration declined as all models increasingly overpredicted risk. However, the random forest and neural network models maintained calibration across ranges of probability, capturing more admissions than did the regression models. The magnitude of overprediction increased over time for the regression models while remaining stable and small for the machine learning models. Changes in the rate of acute kidney injury were strongly linked to increasing overprediction, while changes in predictor-outcome associations corresponded with diverging patterns of calibration drift across methods. CONCLUSIONS: Efficient and effective updating protocols will be essential for maintaining accuracy of, user confidence in, and safety of personalized risk predictions to support decision-making. Model updating protocols should be tailored to account for variations in calibration drift across methods and respond to periods of rapid performance drift rather than be limited to regularly scheduled annual or biannual intervals. Sharon E. Davis, Thomas A. Lasko, Guanhua Chen 0002, Edward D. Siew, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 1 |
| 2016 | Adoption of Secure Messaging in a Patient Portal across Pediatric Specialties
Mary Masterman, Robert M. Cronin, Sharon E. Davis, Jared Shenson, Gretchen Purcell Jackson |
AMIA | 3 |
| 2016 | Use of a Patient Portal During Hospital Admissions to Surgical Services
Jamie R. Robinson, Sharon E. Davis, Robert M. Cronin, Gretchen Purcell Jackson |
AMIA | 2 |
| 2015 | Health Literacy, Education Levels, and Patient Portal Usage During Hospitalizations
Sharon E. Davis, Chandra Y. Osborn, Sunil Kripalani, Kathryn Goggins, Gretchen Purcell Jackson |
AMIA | 1 |