VLDB 2026 Research / reviewers in the wild / expert
Michael J. Pencina
dblp:124/9518
· DBLP profile ↗
20ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-5798-8855ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A federated learning framework for ethical dynamic treatment allocation across heterogeneous hospitals
Xenia Konti, Nicoleta J. Economou-Zavlanos, Yi Shen 0011, Giorgos B. Stamou, Armando Bedoya, Michael J. Pencina, Chuan Hong, Michael M. Zavlanos |
J. Biomed. Informatics | 6 |
| 2025 | Exploring trade-offs in equitable stroke risk prediction with parity-constrained and race-free models
Matthew Engelhard, Daniel Wojdyla, Michael J. Pencina, Ricardo Henao |
Artif. Intell. Medicine | 4 |
| 2025 | Application of unified health large language model evaluation framework to In-Basket message replies: bridging qualitative and quantitative assessmentsabstractOBJECTIVES: Large language models (LLMs) are increasingly utilized in healthcare, transforming medical practice through advanced language processing capabilities. However, the evaluation of LLMs predominantly relies on human qualitative assessment, which is time-consuming, resource-intensive, and may be subject to variability and bias. There is a pressing need for quantitative metrics to enable scalable, objective, and efficient evaluation. MATERIALS AND METHODS: We propose a unified evaluation framework that bridges qualitative and quantitative methods to assess LLM performance in healthcare settings. This framework maps evaluation aspects-such as linguistic quality, efficiency, content integrity, trustworthiness, and usefulness-to both qualitative assessments and quantitative metrics. We apply our approach to empirically evaluate the Epic In-Basket feature, which uses LLM to generate patient message replies. RESULTS: The empirical evaluation demonstrates that while Artificial Intelligence (AI)-generated replies exhibit high fluency, clarity, and minimal toxicity, they face challenges with coherence and completeness. Clinicians' manual decision to use AI-generated drafts correlates strongly with quantitative metrics, suggesting that quantitative metrics have the potential to reduce human effort in the evaluation process and make it more scalable. DISCUSSION: Our study highlights the potential of a unified evaluation framework that integrates qualitative and quantitative methods, enabling scalable and systematic assessments of LLMs in healthcare. Automated metrics streamline evaluation and monitoring processes, but their effective use depends on alignment with human judgment, particularly for aspects requiring contextual interpretation. As LLM applications expand, refining evaluation strategies and fostering interdisciplinary collaboration will be critical to maintaining high standards of accuracy, ethics, and regulatory compliance. CONCLUSION: Our unified evaluation framework bridges the gap between qualitative human assessments and automated quantitative metrics, enhancing the reliability and scalability of LLM evaluations in healthcare. While automated quantitative evaluations are not ready to fully replace qualitative human evaluations, they can be used to enhance the process and, with relevant benchmarks derived from the unified framework proposed here, they can be applied to LLM monitoring and evaluation of updated versions of the original technology evaluated using qualitative human standards. Chuan Hong, Anand Chowdhury, Anthony D. Sorrentino, Monica Agrawal, Armando Bedoya, Sophia Bessias, Nicoleta J. Economou-Zavlanos, Ian Wong, Christian Pean, Kathryn I. Pollak, Eric G. Poon, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 14 |
| 2025 | Enabling inclusive systematic reviews: incorporating preprint articles with large language model-driven evaluationsabstractOBJECTIVES: Systematic reviews in comparative effectiveness research require timely evidence synthesis. With the rapid advancement of medical research, preprint articles play an increasingly important role in accelerating knowledge dissemination. However, as preprint articles are not peer-reviewed before publication, their quality varies significantly, posing challenges for evidence inclusion in systematic reviews. MATERIALS AND METHODS: We developed AutoConfidenceScore (automated confidence score assessment), an advanced framework for predicting preprint publication, which reduces reliance on manual curation and expands the range of predictors, including three key advancements: (1) automated data extraction using natural language processing techniques, (2) semantic embeddings of titles and abstracts, and (3) large language model (LLM)-driven evaluation scores. Additionally, we employed two prediction models: a random forest classifier for binary outcome and a survival cure model that predicts both binary outcome and publication risk over time. RESULTS: The random forest classifier achieved an area under the receiver operating characteristic curve (AUROC) of 0.747 using all features. The survival cure model achieved an AUROC of 0.731 for binary outcome prediction and a concordance index of 0.667 for time-to-publication risk. DISCUSSION: Our study advances the framework for preprint publication prediction through automated data extraction and multiple feature integration. By combining semantic embeddings with LLM-driven evaluations, AutoConfidenceScore significantly enhances predictive performance while reducing manual annotation burden. CONCLUSION: AutoConfidenceScore has the potential to facilitate incorporation of preprint articles during the appraisal phase of systematic reviews, supporting researchers in more effective utilization of preprint resources. Rui Yang 0016, Jiayi Tong, Nan Liu 0003, Christopher J. Lindsell, Michael J. Pencina, Yong Chen 0016, Chuan Hong |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Adaptive Discretization for Event PredicTion (ADEPT)abstractRecently developed survival analysis methods improve upon existing approaches by predicting the probability of event occurrence in each of a number pre-specified (discrete) time intervals. By avoiding placing strong parametric assumptions on the event density, this approach tends to improve prediction performance, particularly when data are plentiful. However, in clinical settings with limited available data, it is often preferable to judiciously partition the event time space into a limited number of intervals well suited to the prediction task at hand. In this work, we develop Adaptive Discretization for Event PredicTion (ADEPT) to learn from data a set of cut points defining such a partition. We show that in two simulated datasets, we are able to recover intervals that match the underlying generative model. We then demonstrate improved prediction performance on three real-world observational datasets, including a large, newly harmonized stroke risk prediction dataset. Finally, we argue that our approach facilitates clinical decision-making by suggesting time intervals that are most appropriate for each task, in the sense that they facilitate more accurate risk prediction. Jimmy Hickey, Ricardo Henao, Daniel Wojdyla, Michael J. Pencina, Matthew Engelhard |
AISTATS | 4 |
| 2024 | Translating ethical and quality principles for the effective, safe and fair development, deployment and use of artificial intelligence technologies in healthcareabstractOBJECTIVE: The complexity and rapid pace of development of algorithmic technologies pose challenges for their regulation and oversight in healthcare settings. We sought to improve our institution's approach to evaluation and governance of algorithmic technologies used in clinical care and operations by creating an Implementation Guide that standardizes evaluation criteria so that local oversight is performed in an objective fashion. MATERIALS AND METHODS: Building on a framework that applies key ethical and quality principles (clinical value and safety, fairness and equity, usability and adoption, transparency and accountability, and regulatory compliance), we created concrete guidelines for evaluating algorithmic technologies at our institution. RESULTS: An Implementation Guide articulates evaluation criteria used during review of algorithmic technologies and details what evidence supports the implementation of ethical and quality principles for trustworthy health AI. Application of the processes described in the Implementation Guide can lead to algorithms that are safer as well as more effective, fair, and equitable upon implementation, as illustrated through 4 examples of technologies at different phases of the algorithmic lifecycle that underwent evaluation at our academic medical center. DISCUSSION: By providing clear descriptions/definitions of evaluation criteria and embedding them within standardized processes, we streamlined oversight processes and educated communities using and developing algorithmic technologies within our institution. CONCLUSIONS: We developed a scalable, adaptable framework for translating principles into evaluation criteria and specific requirements that support trustworthy implementation of algorithmic technologies in patient care and healthcare operations. Nicoleta J. Economou-Zavlanos, Sophia Bessias, Michael P. Cary, Armando Bedoya, Benjamin Goldstein 0001, John Eric Jelovsek, Cara O'Brien, Nancy Walden, Matthew Elmore, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Michael E. Lipkin, Afreen Idris Shariff, Michael Gao, David Leverenz, Ricardo Henao, David Y. Ming, David M. Gallagher, Michael J. Pencina, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 21 |
| 2024 | Trans-Balance: Reducing demographic disparity for prediction models in the presence of class imbalance
Chuan Hong, Molei Liu, Daniel Wojdyla, Jimmy Hickey, Michael J. Pencina, Ricardo Henao |
J. Biomed. Informatics | 5 |
| 2023 | Semi-supervised calibration of noisy event risk (SCANER) with electronic health records
Chuan Hong, Qianyu Yuan, Kelly Cho, Katherine P. Liao, Michael J. Pencina, David C. Christiani, Tianxi Cai |
J. Biomed. Informatics | 6 |
| 2023 | Calibration and Uncertainty in Neural Time-to-Event ModelingabstractModels for predicting the time of a future event are crucial for risk assessment, across a diverse range of applications. Existing time-to-event (survival) models have focused primarily on preserving pairwise ordering of estimated event times (i.e., relative risk). We propose neural time-to-event models that account for calibration and uncertainty while predicting accurate absolute event times. Specifically, an adversarial nonparametric model is introduced for estimating matched time-to-event distributions for probabilistically concentrated and accurate predictions. We also consider replacing the discriminator of the adversarial nonparametric model with a survival-function matching estimator that accounts for model calibration. The proposed estimator can be used as a means of estimating and comparing conditional survival distributions while accounting for the predictive uncertainty of probabilistic models. Extensive experiments show that the distribution matching methods outperform existing approaches in terms of both calibration and concentration of time-to-event distributions. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Karen Chandross, Michael J. Pencina, Lawrence Carin, Ricardo Henao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | A framework for the oversight and local deployment of safe and high-quality prediction modelsabstractArtificial intelligence/machine learning models are being rapidly developed and used in clinical practice. However, many models are deployed without a clear understanding of clinical or operational impact and frequently lack monitoring plans that can detect potential safety signals. There is a lack of consensus in establishing governance to deploy, pilot, and monitor algorithms within operational healthcare delivery workflows. Here, we describe a governance framework that combines current regulatory best practices and lifecycle management of predictive models being used for clinical care. Since January 2021, we have successfully added models to our governance portfolio and are currently managing 52 models. Armando Bedoya, Nicoleta J. Economou-Zavlanos, Benjamin Goldstein 0001, Allison Young, John Eric Jelovsek, Cara O'Brien, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Erich Huang, Eric G. Poon, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 13 |
| 2022 | Observability and its impact on differential bias for clinical prediction modelsabstractOBJECTIVE: Electronic health records have incomplete capture of patient outcomes. We consider the case when observability is differential across a predictor. Including such a predictor (sensitive variable) can lead to algorithmic bias, potentially exacerbating health inequities. MATERIALS AND METHODS: We define bias for a clinical prediction model (CPM) as the difference between the true and estimated risk, and differential bias as bias that differs across a sensitive variable. We illustrate the genesis of differential bias via a 2-stage process, where conditional on having the outcome of interest, the outcome is differentially observed. We use simulations and a real-data example to demonstrate the possible impact of including a sensitive variable in a CPM. RESULTS: If there is differential observability based on a sensitive variable, including it in a CPM can induce differential bias. However, if the sensitive variable impacts the outcome but not observability, it is better to include it. When a sensitive variable impacts both observability and the outcome no simple recommendation can be provided. We show that one cannot use observed data to detect differential bias. DISCUSSION: Our study furthers the literature on observability, showing that differential observability can lead to algorithmic bias. This highlights the importance of considering whether to include sensitive variables in CPMs. CONCLUSION: Including a sensitive variable in a CPM depends on whether it truly affects the outcome or just the observability of the outcome. Since this cannot be distinguished with observed data, observability is an implicit assumption of CPMs. Mengying Yan, Michael J. Pencina, L. Ebony Boulware, Benjamin Goldstein 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Understanding Algorithmic Bias in Clinical Prediction Models
Mengying Yan, Michael J. Pencina, Benjamin Goldstein 0001 |
AMIA | 2 |
| 2019 | Informatics-Enabled Learning Health Systems: Strategies for Success from Four Academic Medical Centers
Eric G. Poon, Charles P. Friedman, Philip R. O. Payne, Michael J. Pencina, Kevin B. Johnson |
AMIA | 4 |
| 2019 | An outcome model approach to transporting a randomized controlled trial results to a target populationabstractOBJECTIVE: Participants enrolled into randomized controlled trials (RCTs) often do not reflect real-world populations. Previous research in how best to transport RCT results to target populations has focused on weighting RCT data to look like the target data. Simulation work, however, has suggested that an outcome model approach may be preferable. Here, we describe such an approach using source data from the 2 × 2 factorial NAVIGATOR (Nateglinide And Valsartan in Impaired Glucose Tolerance Outcomes Research) trial, which evaluated the impact of valsartan and nateglinide on cardiovascular outcomes and new-onset diabetes in a prediabetic population. MATERIALS AND METHODS: Our target data consisted of people with prediabetes serviced at the Duke University Health System. We used random survival forests to develop separate outcome models for each of the 4 treatments, estimating the 5-year risk difference for progression to diabetes, and estimated the treatment effect in our local patient populations, as well as subpopulations, and compared the results with the traditional weighting approach. RESULTS: Our models suggested that the treatment effect for valsartan in our patient population was the same as in the trial, whereas for nateglinide treatment effect was stronger than observed in the original trial. Our effect estimates were more efficient than the weighting approach and we effectively estimated subgroup differences. CONCLUSIONS: The described method represents a straightforward approach to efficiently transporting an RCT result to any target population. Benjamin Goldstein 0001, Matthew Phelan, Neha J. Pagidipati, Rury R. Holman, Michael J. Pencina, Elizabeth A. Stuart |
J. Am. Medical Informatics Assoc. | 5 |
| 2018 | The Duke Health Data Science Internship Program: Integrating the Educational Mission into Real-World Research
Shelley A. Rusincovitch, Lisa Wruck, Ricardo Henao, Larisa Rodgers, Allison Dunning, Peter Merrill, Hillary Mulder, Robert Overton, Matthew Phelan, Erich Huang, Lawrence Carin, Michael J. Pencina |
AMIA | 12 |
| 2018 | Microsimulation model to predict incremental value of biomarkers added to prognostic modelsabstractIt is unclear to what extent simulated versions of real data can be used to assess potential value of new biomarkers added to prognostic risk models. Using data on 4522 women and 3969 men who contributed information to the Framingham CVD risk prediction tool, we develop a simulation model that allows assessment of the added contribution of new biomarkers. The simulated model matches closely the one obtained using real data: discrimination area under the curve (AUC) on simulated vs actual data is 0.800 vs 0.799 in women and 0.778 vs 0.776 in men. Positive correlation with standard risk factors decreases the impact of new biomarkers (ΔAUC 0.002-0.024), but negative correlation leads to stronger effects (ΔAUC 0.026-0.101) than no correlation (ΔAUC 0.003-0.051). We suggest that researchers construct simulation models similar to the one proposed here before embarking on larger, expensive biomarker studies based on actual data. Karol M. Pencina, Ralph B. D'Agostino, Ramachandran S. Vasan, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 4 |
| 2017 | Developing a framework for a comprehensive data sharing program
Asba Tasneem, Karen Chiswell, Brian McCourt, Matt Gross, Eric D. Peterson, Michael J. Pencina |
AMIA | 6 |
| 2017 | Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic reviewabstractOBJECTIVE: Electronic health records (EHRs) are an increasingly common data source for clinical risk prediction, presenting both unique analytic opportunities and challenges. We sought to evaluate the current state of EHR based risk prediction modeling through a systematic review of clinical prediction studies using EHR data. METHODS: We searched PubMed for articles that reported on the use of an EHR to develop a risk prediction model from 2009 to 2014. Articles were extracted by two reviewers, and we abstracted information on study design, use of EHR data, model building, and performance from each publication and supplementary documentation. RESULTS: We identified 107 articles from 15 different countries. Studies were generally very large (median sample size = 26 100) and utilized a diverse array of predictors. Most used validation techniques (n = 94 of 107) and reported model coefficients for reproducibility (n = 83). However, studies did not fully leverage the breadth of EHR data, as they uncommonly used longitudinal information (n = 37) and employed relatively few predictor variables (median = 27 variables). Less than half of the studies were multicenter (n = 50) and only 26 performed validation across sites. Many studies did not fully address biases of EHR data such as missing data or loss to follow-up. Average c-statistics for different outcomes were: mortality (0.84), clinical prediction (0.83), hospitalization (0.71), and service utilization (0.71). CONCLUSIONS: EHR data present both opportunities and challenges for clinical risk prediction. There is room for improvement in designing such studies. Benjamin Goldstein 0001, Ann Marie Navar, Michael J. Pencina, John P. A. Ioannidis |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | Predicting mortality over different time horizons: which data elements are needed?abstractOBJECTIVE: Electronic health records (EHRs) are a resource for "big data" analytics, containing a variety of data elements. We investigate how different categories of information contribute to prediction of mortality over different time horizons among patients undergoing hemodialysis treatment. MATERIAL AND METHODS: We derived prediction models for mortality over 7 time horizons using EHR data on older patients from a national chain of dialysis clinics linked with administrative data using LASSO (least absolute shrinkage and selection operator) regression. We assessed how different categories of information relate to risk assessment and compared discrete models to time-to-event models. RESULTS: The best predictors used all the available data (c-statistic ranged from 0.72-0.76), with stronger models in the near term. While different variable groups showed different utility, exclusion of any particular group did not lead to a meaningfully different risk assessment. Discrete time models performed better than time-to-event models. CONCLUSIONS: Different variable groups were predictive over different time horizons, with vital signs most predictive for near-term mortality and demographic and comorbidities more important in long-term mortality. Benjamin Goldstein 0001, Michael J. Pencina, Maria E. Montez-Rath, Wolfgang C. Winkelmayer |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Assessing electronic health record phenotypes against gold-standard diagnostic criteria for diabetes mellitusabstractOBJECTIVE: We assessed the sensitivity and specificity of 8 electronic health record (EHR)-based phenotypes for diabetes mellitus against gold-standard American Diabetes Association (ADA) diagnostic criteria via chart review by clinical experts. MATERIALS AND METHODS: We identified EHR-based diabetes phenotype definitions that were developed for various purposes by a variety of users, including academic medical centers, Medicare, the New York City Health Department, and pharmacy benefit managers. We applied these definitions to a sample of 173 503 patients with records in the Duke Health System Enterprise Data Warehouse and at least 1 visit over a 5-year period (2007-2011). Of these patients, 22 679 (13%) met the criteria of 1 or more of the selected diabetes phenotype definitions. A statistically balanced sample of these patients was selected for chart review by clinical experts to determine the presence or absence of type 2 diabetes in the sample. RESULTS: The sensitivity (62-94%) and specificity (95-99%) of EHR-based type 2 diabetes phenotypes (compared with the gold standard ADA criteria via chart review) varied depending on the component criteria and timing of observations and measurements. DISCUSSION AND CONCLUSIONS: Researchers using EHR-based phenotype definitions should clearly specify the characteristics that comprise the definition, variations of ADA criteria, and how different phenotype definitions and components impact the patient populations retrieved and the intended application. Careful attention to phenotype definitions is critical if the promise of leveraging EHR data to improve individual and population health is to be fulfilled. Susan E. Spratt, Katherine Pereira, Bradi B. Granger, Bryan C. Batch, Matthew Phelan, Michael J. Pencina, Marie Lynn Miranda, L. Ebony Boulware, Joseph E. Lucas, Charlotte L. Nelson, Benjamin Neely, Benjamin Goldstein 0001, Pamela Barth, Rachel L. Richesson, Isaretta L. Riley, Leonor Corsino, Eugenia R. McPeek Hinz, Shelley A. Rusincovitch, Jennifer Green, Anna Beth Barton, Carly Kelley, Kristen Hyland, Monica Tang, Amanda Elliott, Ewa Ruel, Alexander Clark, Melanie Mabrey, Kay Lyn Morrissey, Jyothi Rao, Beatrice Hong, Marjorie Pierre-Louis, Katherine Kelly, Nicole E. Jelesoff |
J. Am. Medical Informatics Assoc. | 6 |