EDBT 2026 Demo / reviewers in the wild / expert
Benjamin Goldstein 0001
dblp:294/8298 · also Benjamin Alan Goldstein
· DBLP profile ↗
26ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-5261-3632ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comparing ambient scribes: a randomized crossover clinical trial addressing ambient scribe technologies' impact on physician burnoutabstractOBJECTIVE: This study aims to compare the effectiveness of 2 ambient AI scribe technologies in reducing physician burnout, improving workflow satisfaction, and enhancing documentation efficiency through a randomized crossover trial. MATERIALS AND METHODS: An open-label randomized crossover trial involving 160 outpatient clinicians was conducted at a tertiary academic medical center. Volunteers were randomized to 2 groups of 80 with 2 crossover periods. We assessed workflow satisfaction (1-7 scale), burnout (Copenhagen Burnout Index), and efficiency metrics (eg, electronic health record time outside scheduled hours, documentation time, etc.). Data was analyzed using Wilcoxon signed-rank tests and generalized linear mixed models. RESULTS: Surveys from 136 respondents were analyzed. Clinicians reported greater improvements in satisfaction with product B (2.51 points on a 7-point scale) compared to product A (1.91 points; mean difference: 0.60, 95% CI: 0.32-0.90). Both tools reduced personal and work burnout scores, but differences between tools were not meaningful. Product B demonstrated greater reductions in average minutes-in-notes per day compared to product A (B - A = -3.19 minutes; 95% CI -4.87 to -1.50). No meaningful differences were observed in pajama time or patient-related burnout. DISCUSSION: Both tools improved workflow satisfaction and reduced burnout, with product B showing superior performance in satisfaction and documentation time. However, efficiency metrics like pajama time were largely unaffected, potentially due to participant selection bias and the study period's timing. CONCLUSION: Product B yielded greater satisfaction and time savings compared to product A, though both tools effectively reduced physician burnout and improved workflow satisfaction. Anand Chowdhury, Michele Casey, Jonathan Wilson, Kathryn I. Pollak, Benjamin Goldstein 0001, Armando Bedoya, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 5 |
| 2026 | Discrete Time Neural Network Models to Address Time-Varying Predictor Importance: An Illustration in Predicting Mortality Over Different Time HorizonsabstractClinical predictive models (CPMs) are crucial for forecasting patient outcomes using available electronic health record (EHR) data. Traditional time-to-event (TTE) models, like the Cox proportional hazards model, assume that hazard ratios remain constant over time, which may not hold in many clinical settings. In this study, we introduce a Discrete Time Neural Network (DTNN) to address these limitations by modeling time-varying predictor importance. The DTNN combines the flexibility of classification models with the advantage of handling time-to-event data, providing a single model fit across multiple time horizons. Using data from patients with end-stage kidney disease (ESKD) undergoing hemodialysis, we demonstrate that the DTNN can flexibly adjust for different risk factors across short-term and long-term mortality predictions. The model was evaluated using cumulative-dynamic area under the receiver operating characteristic (CD-AUROC) to compare patients who remained event-free up to a given time t with those who experienced the event before t. Our results show that the DTNN outperforms traditional TTE models and provides robust predictions across varying time intervals, making it an appealing choice for clinical settings where time-varying predictor importance is essential. Mufan Wang, Matthew Engelhard, Patrick H. Pun, Benjamin Goldstein 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Machine learning-based prediction models in medical decision-making in kidney disease: patient, caregiver, and clinician perspectives on trust and appropriate useabstractOBJECTIVES: This study aims to improve the ethical use of machine learning (ML)-based clinical prediction models (CPMs) in shared decision-making for patients with kidney failure on dialysis. We explore factors that inform acceptability, interpretability, and implementation of ML-based CPMs among multiple constituent groups. MATERIALS AND METHODS: We collected and analyzed qualitative data from focus groups with varied end users, including: dialysis support providers (clinical providers and additional dialysis support providers such as dialysis clinic staff and social workers); patients; patients' caregivers (n = 52). RESULTS: Participants were broadly accepting of ML-based CPMs, but with concerns on data sources, factors included in the model, and accuracy. Use was desired in conjunction with providers' views and explanations. Differences among respondent types were minimal overall but most prevalent in discussions of CPM presentation and model use. DISCUSSION AND CONCLUSION: Evidence of acceptability of ML-based CPM usage provides support for ethical use, but numerous specific considerations in acceptability, model construction, and model use for shared clinical decision-making must be considered. There are specific steps that could be taken by data scientists and health systems to engender use that is accepted by end users and facilitates trust, but there are also ongoing barriers or challenges in addressing desires for use. This study contributes to emerging literature on interpretability, mechanisms for sharing complexities, including uncertainty regarding the model results, and implications for decision-making. It examines numerous stakeholder groups including providers, patients, and caregivers to provide specific considerations that can influence health system use and provide a basis for future research. Jessica Sperling, Whitney Welsh, Erin Haseley, Stella Quenstedt, Perusi B. Muhigaba, Adrian Brown, Patti Ephraim, Tariq Shafi, Michael Waitzkin, David J. Casarett, Benjamin Goldstein 0001 |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Contrastive Learning for Clinical Outcome Prediction with Partial Data SourcesabstractThe use of machine learning models to predict clinical outcomes from (longitudinal) electronic health record (EHR) data is becoming increasingly popular due to advances in deep architectures, representation learning, and the growing availability of large EHR datasets. Existing models generally assume access to the same data sources during both training and inference stages. However, this assumption is often challenged by the fact that real-world clinical datasets originate from various data sources (with distinct sets of covariates), which though can be available for training (in a research or retrospective setting), are more realistically only partially available (a subset of such sets) for inference when deployed. So motivated, we introduce Contrastive Learning for clinical Outcome Prediction with Partial data Sources (CLOPPS), that trains encoders to capture information across different data sources and then leverages them to build classifiers restricting access to a single data source. This approach can be used with existing cross-sectional or longitudinal outcome classification models. We present experiments on two real-world datasets demonstrating that CLOPPS consistently outperforms strong baselines in several practical scenarios. Jonathan Wilson, Benjamin Goldstein 0001, Ricardo Henao |
ICML | 3 |
| 2024 | Translating ethical and quality principles for the effective, safe and fair development, deployment and use of artificial intelligence technologies in healthcareabstractOBJECTIVE: The complexity and rapid pace of development of algorithmic technologies pose challenges for their regulation and oversight in healthcare settings. We sought to improve our institution's approach to evaluation and governance of algorithmic technologies used in clinical care and operations by creating an Implementation Guide that standardizes evaluation criteria so that local oversight is performed in an objective fashion. MATERIALS AND METHODS: Building on a framework that applies key ethical and quality principles (clinical value and safety, fairness and equity, usability and adoption, transparency and accountability, and regulatory compliance), we created concrete guidelines for evaluating algorithmic technologies at our institution. RESULTS: An Implementation Guide articulates evaluation criteria used during review of algorithmic technologies and details what evidence supports the implementation of ethical and quality principles for trustworthy health AI. Application of the processes described in the Implementation Guide can lead to algorithms that are safer as well as more effective, fair, and equitable upon implementation, as illustrated through 4 examples of technologies at different phases of the algorithmic lifecycle that underwent evaluation at our academic medical center. DISCUSSION: By providing clear descriptions/definitions of evaluation criteria and embedding them within standardized processes, we streamlined oversight processes and educated communities using and developing algorithmic technologies within our institution. CONCLUSIONS: We developed a scalable, adaptable framework for translating principles into evaluation criteria and specific requirements that support trustworthy implementation of algorithmic technologies in patient care and healthcare operations. Nicoleta J. Economou-Zavlanos, Sophia Bessias, Michael P. Cary, Armando Bedoya, Benjamin Goldstein 0001, John Eric Jelovsek, Cara O'Brien, Nancy Walden, Matthew Elmore, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Michael E. Lipkin, Afreen Idris Shariff, Michael Gao, David Leverenz, Ricardo Henao, David Y. Ming, David M. Gallagher, Michael J. Pencina, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 5 |
| 2024 | A conditional multi-label model to improve prediction of a rare outcome: An illustration predicting autism diagnosis
Wei A. Huang, Matthew Engelhard, Marika Coffman, Elliot D. Hill, Qin Weng, Abby Scheer, Gary Maslow, Ricardo Henao, Geraldine Dawson, Benjamin Goldstein 0001 |
J. Biomed. Informatics | 10 |
| 2022 | A framework for the oversight and local deployment of safe and high-quality prediction modelsabstractArtificial intelligence/machine learning models are being rapidly developed and used in clinical practice. However, many models are deployed without a clear understanding of clinical or operational impact and frequently lack monitoring plans that can detect potential safety signals. There is a lack of consensus in establishing governance to deploy, pilot, and monitor algorithms within operational healthcare delivery workflows. Here, we describe a governance framework that combines current regulatory best practices and lifecycle management of predictive models being used for clinical care. Since January 2021, we have successfully added models to our governance portfolio and are currently managing 52 models. Armando Bedoya, Nicoleta J. Economou-Zavlanos, Benjamin Goldstein 0001, Allison Young, John Eric Jelovsek, Cara O'Brien, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Erich Huang, Eric G. Poon, Michael J. Pencina |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | Observability and its impact on differential bias for clinical prediction modelsabstractOBJECTIVE: Electronic health records have incomplete capture of patient outcomes. We consider the case when observability is differential across a predictor. Including such a predictor (sensitive variable) can lead to algorithmic bias, potentially exacerbating health inequities. MATERIALS AND METHODS: We define bias for a clinical prediction model (CPM) as the difference between the true and estimated risk, and differential bias as bias that differs across a sensitive variable. We illustrate the genesis of differential bias via a 2-stage process, where conditional on having the outcome of interest, the outcome is differentially observed. We use simulations and a real-data example to demonstrate the possible impact of including a sensitive variable in a CPM. RESULTS: If there is differential observability based on a sensitive variable, including it in a CPM can induce differential bias. However, if the sensitive variable impacts the outcome but not observability, it is better to include it. When a sensitive variable impacts both observability and the outcome no simple recommendation can be provided. We show that one cannot use observed data to detect differential bias. DISCUSSION: Our study furthers the literature on observability, showing that differential observability can lead to algorithmic bias. This highlights the importance of considering whether to include sensitive variables in CPMs. CONCLUSION: Including a sensitive variable in a CPM depends on whether it truly affects the outcome or just the observability of the outcome. Since this cannot be distinguished with observed data, observability is an implicit assumption of CPMs. Mengying Yan, Michael J. Pencina, L. Ebony Boulware, Benjamin Goldstein 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | AutoScore-Survival: Developing interpretable machine learning-based time-to-event scores with right-censored survival data
Feng Xie 0004, Yilin Ning, Benjamin Goldstein 0001, Marcus Eng Hock Ong, Nan Liu 0003, Bibhas Chakraborty |
J. Biomed. Informatics | 4 |
| 2022 | AutoScore-Imbalance: An interpretable machine learning tool for development of clinical scores with rare events data
Feng Xie 0004, Marcus Eng Hock Ong, Yilin Ning, Marcel Lucas Chee, Seyed Ehsan Saffari, Hairil Rizal Abdullah, Benjamin Goldstein 0001, Bibhas Chakraborty, Nan Liu 0003 |
J. Biomed. Informatics | 8 |
| 2021 | Variational Disentanglement for Rare Event ModelingabstractCombining the increasing availability and abundance of healthcare data and the current advances in machine learning methods have created renewed opportunities to improve clinical decision support systems. However, in healthcare risk prediction applications, the proportion of cases with the condition (label) of interest is often very low relative to the available sample size. Though very prevalent in healthcare, such imbalanced classification settings are also common and challenging in many other scenarios. So motivated, we propose a variational disentanglement approach to semi-parametrically learn from rare events in heavily imbalanced classification problems. Specifically, we leverage the imposed extreme-distribution behavior on a latent space to extract information from low-prevalence events, and develop a robust prediction arm that joins the merits of the generalized additive model and isotonic neural nets. Results on synthetic studies and diverse real-world datasets, including mortality prediction on a COVID-19 cohort, demonstrate that the proposed approach outperforms existing alternatives. Zidi Xiu, Chenyang Tao, Michael Gao, Connor Davis, Benjamin Goldstein 0001, Ricardo Henao |
AAAI | 5 |
| 2021 | Comparison of a patient cohort and predictive models derived from local academic medical centers versus a national health database
Courtney Page, Conrad Sweitek, Cliona Molony, Karen Chandross, Benjamin Goldstein 0001 |
AMIA | 5 |
| 2021 | Understanding Algorithmic Bias in Clinical Prediction Models
Mengying Yan, Michael J. Pencina, Benjamin Goldstein 0001 |
AMIA | 3 |
| 2021 | Supercharging Imbalanced Data Learning With Energy-based Contrastive Representation TransferabstractDealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest.In computer vision and NLP, learning from datasets with long-tail behavior is a recurring theme, especially for naturally occurring labels. Existing solutions mostly appeal to sampling or weighting adjustments to alleviate the extreme imbalance, or impose inductive bias to prioritize generalizable associations. Here we take a novel perspective to promote sample efficiency and model generalization based on the invariance principles of causality. Our contribution posits a meta-distributional scenario, where the causal generating mechanism for label-conditional features is invariant across different labels. Such causal assumption enables efficient knowledge transfer from the dominant classes to their under-represented counterparts, even if their feature distributions show apparent disparities. This allows us to leverage a causal data augmentation procedure to enlarge the representation of minority classes. Our development is orthogonal to the existing imbalanced data learning techniques thus can be seamlessly integrated. The proposed approach is validated on an extensive set of synthetic and real-world tasks against state-of-the-art solutions. Junya Chen, Zidi Xiu, Benjamin Goldstein 0001, Ricardo Henao, Lawrence Carin, Chenyang Tao |
NeurIPS | 3 |
| 2020 | Identified themes of interactive visualizations overlayed onto EHR data: an example of improving birth center operating room efficiencyabstractOBJECTIVE: While electronic health record (EHR) systems store copious amounts of patient data, aggregating those data across patients can be challenging. Visual analytic tools that integrate with EHR systems allow clinicians to gain better insight and understanding into clinical care and management. We report on our experience building Tableau-based visualizations and integrating them into our EHR system. MATERIALS AND METHODS: Visual analytic tools were created as part of 12 clinician-initiated quality improvement projects. We built the visual analytic tools in Tableau and linked it within our EPIC environment. We identified 5 visual themes that spanned the various projects. To illustrate these themes, we choose 1 exemplary project which aimed to improve obstetric operating room efficiency. RESULTS: Across our 12 projects, we identified 5 visual themes that are integral to project success: scheduling & optimization (in 11/12 projects); provider assessment (10/12); executive assessment (8/12); patient outcomes (7/12); and control and goal charts (2/12). DISCUSSION: Many visualizations share common themes. Identification of these themes has allowed our internal team to be more efficient and directed in developing visualizations for future projects. CONCLUSION: Organizing visual analytics into themes can allow informatics teams to more efficiently provide visual products to clinical collaborators. Andrew Stirling, Tracy Tubb, Emily S. Reiff, Chad A. Grotegut, Jennifer Gagnon, Gail Bradley, Eric G. Poon, Benjamin Goldstein 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2019 | An outcome model approach to transporting a randomized controlled trial results to a target populationabstractOBJECTIVE: Participants enrolled into randomized controlled trials (RCTs) often do not reflect real-world populations. Previous research in how best to transport RCT results to target populations has focused on weighting RCT data to look like the target data. Simulation work, however, has suggested that an outcome model approach may be preferable. Here, we describe such an approach using source data from the 2 × 2 factorial NAVIGATOR (Nateglinide And Valsartan in Impaired Glucose Tolerance Outcomes Research) trial, which evaluated the impact of valsartan and nateglinide on cardiovascular outcomes and new-onset diabetes in a prediabetic population. MATERIALS AND METHODS: Our target data consisted of people with prediabetes serviced at the Duke University Health System. We used random survival forests to develop separate outcome models for each of the 4 treatments, estimating the 5-year risk difference for progression to diabetes, and estimated the treatment effect in our local patient populations, as well as subpopulations, and compared the results with the traditional weighting approach. RESULTS: Our models suggested that the treatment effect for valsartan in our patient population was the same as in the trial, whereas for nateglinide treatment effect was stronger than observed in the original trial. Our effect estimates were more efficient than the weighting approach and we effectively estimated subgroup differences. CONCLUSIONS: The described method represents a straightforward approach to efficiently transporting an RCT result to any target population. Benjamin Goldstein 0001, Matthew Phelan, Neha J. Pagidipati, Rury R. Holman, Michael J. Pencina, Elizabeth A. Stuart |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | How and when informative visit processes can bias inference when using electronic health records data for clinical researchabstractOBJECTIVE: Electronic health records (EHR) data have become a central data source for clinical research. One concern for using EHR data is that the process through which individuals engage with the health system, and find themselves within EHR data, can be informative. We have termed this process informed presence. In this study we use simulation and real data to assess how the informed presence can impact inference. MATERIALS AND METHODS: We first simulated a visit process where a series of biomarkers were observed informatively and uninformatively over time. We further compared inference derived from a randomized control trial (ie, uninformative visits) and EHR data (ie, potentially informative visits). RESULTS: We find that only when there is both a strong association between the biomarker and the outcome as well as the biomarker and the visit process is there bias. Moreover, once there are some uninformative visits this bias is mitigated. In the data example we find, that when the "true" associations are null, there is no observed bias. DISCUSSION: These results suggest that an informative visit process can exaggerate an association but cannot induce one. Furthermore, careful study design can, mitigate the potential bias when some noninformative visits are included. CONCLUSIONS: While there are legitimate concerns regarding biases that "messy" EHR data may induce, the conditions for such biases are extreme and can be accounted for. Benjamin Goldstein 0001, Matthew Phelan, Neha J. Pagidipati, Sarah B. Peskoe |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Lessons Learned from an EHR-Based Population Health Datamart: Southeastern Diabetes Initiative (SEDI)
Ursula Rogers, Shelley A. Rusincovitch, Matthew Phelan, Nigel B. Neely, Benjamin Goldstein 0001 |
AMIA | 5 |
| 2018 | Adversarial Time-to-Event ModelingabstractModern health data science applications leverage abundant molecular and electronic health data, providing opportunities for machine learning to build statistical models to support clinical practice. Time-to-event analysis, also called survival analysis, stands as one of the most representative examples of such statistical models. We present a deep-network-based approach that leverages adversarial learning to address a key challenge in modern time-to-event modeling: nonparametric estimation of event-time distributions. We also introduce a principled cost function to exploit information from censored events (events that occur subsequent to the observation window). Unlike most time-to-event models, we focus on the estimation of time-to-event distributions, rather than time ordering. We validate our model on both benchmark and real datasets, demonstrating that the proposed formulation yields significant performance gains relative to a parametric alternative, which we also propose. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Courtney Page, Benjamin Goldstein 0001, Lawrence Carin, Ricardo Henao |
ICML | 5 |
| 2018 | Designing risk prediction models for ambulatory no-shows across different specialties and clinicsabstractObjective: As available data increases, so does the opportunity to develop risk scores on more refined patient populations. In this paper we assessed the ability to derive a risk score for a patient no-showing to a clinic visit. Methods: Using data from 2 264 235 outpatient appointments we assessed the performance of models built across 14 different specialties and 55 clinics. We used regularized logistic regression models to fit and assess models built on the health system, specialty, and clinic levels. We evaluated fits based on their discrimination and calibration. Results: Overall, the results suggest that a relatively robust risk score for patient no-shows could be derived with an average C-statistic of 0.83 across clinic level models and strong calibration. Moreover, the clinic specific models, even with lower training set sizes, often performed better than the more general models. Examination of the individual models showed that risk factors had different degrees of predictability across the different specialties. Implementation of optimal modeling strategies would lead to capturing an additional 4819 no-shows per-year. Conclusion: Overall, this work highlights both the opportunity for and the importance of leveraging the available electronic health record data to develop more refined risk models. Xiruo Ding, Ziad Gellad, III Chad Mather, Pamela Barth, Eric G. Poon, Mark Newman, Benjamin Goldstein 0001 |
J. Am. Medical Informatics Assoc. | 7 |
| 2017 | Opportunities and challenges in developing risk prediction models with electronic health records data: a systematic reviewabstractOBJECTIVE: Electronic health records (EHRs) are an increasingly common data source for clinical risk prediction, presenting both unique analytic opportunities and challenges. We sought to evaluate the current state of EHR based risk prediction modeling through a systematic review of clinical prediction studies using EHR data. METHODS: We searched PubMed for articles that reported on the use of an EHR to develop a risk prediction model from 2009 to 2014. Articles were extracted by two reviewers, and we abstracted information on study design, use of EHR data, model building, and performance from each publication and supplementary documentation. RESULTS: We identified 107 articles from 15 different countries. Studies were generally very large (median sample size = 26 100) and utilized a diverse array of predictors. Most used validation techniques (n = 94 of 107) and reported model coefficients for reproducibility (n = 83). However, studies did not fully leverage the breadth of EHR data, as they uncommonly used longitudinal information (n = 37) and employed relatively few predictor variables (median = 27 variables). Less than half of the studies were multicenter (n = 50) and only 26 performed validation across sites. Many studies did not fully address biases of EHR data such as missing data or loss to follow-up. Average c-statistics for different outcomes were: mortality (0.84), clinical prediction (0.83), hospitalization (0.71), and service utilization (0.71). CONCLUSIONS: EHR data present both opportunities and challenges for clinical risk prediction. There is room for improvement in designing such studies. Benjamin Goldstein 0001, Ann Marie Navar, Michael J. Pencina, John P. A. Ioannidis |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Predicting mortality over different time horizons: which data elements are needed?abstractOBJECTIVE: Electronic health records (EHRs) are a resource for "big data" analytics, containing a variety of data elements. We investigate how different categories of information contribute to prediction of mortality over different time horizons among patients undergoing hemodialysis treatment. MATERIAL AND METHODS: We derived prediction models for mortality over 7 time horizons using EHR data on older patients from a national chain of dialysis clinics linked with administrative data using LASSO (least absolute shrinkage and selection operator) regression. We assessed how different categories of information relate to risk assessment and compared discrete models to time-to-event models. RESULTS: The best predictors used all the available data (c-statistic ranged from 0.72-0.76), with stronger models in the near term. While different variable groups showed different utility, exclusion of any particular group did not lead to a meaningfully different risk assessment. Discrete time models performed better than time-to-event models. CONCLUSIONS: Different variable groups were predictive over different time horizons, with vital signs most predictive for near-term mortality and demographic and comorbidities more important in long-term mortality. Benjamin Goldstein 0001, Michael J. Pencina, Maria E. Montez-Rath, Wolfgang C. Winkelmayer |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Assessing electronic health record phenotypes against gold-standard diagnostic criteria for diabetes mellitusabstractOBJECTIVE: We assessed the sensitivity and specificity of 8 electronic health record (EHR)-based phenotypes for diabetes mellitus against gold-standard American Diabetes Association (ADA) diagnostic criteria via chart review by clinical experts. MATERIALS AND METHODS: We identified EHR-based diabetes phenotype definitions that were developed for various purposes by a variety of users, including academic medical centers, Medicare, the New York City Health Department, and pharmacy benefit managers. We applied these definitions to a sample of 173 503 patients with records in the Duke Health System Enterprise Data Warehouse and at least 1 visit over a 5-year period (2007-2011). Of these patients, 22 679 (13%) met the criteria of 1 or more of the selected diabetes phenotype definitions. A statistically balanced sample of these patients was selected for chart review by clinical experts to determine the presence or absence of type 2 diabetes in the sample. RESULTS: The sensitivity (62-94%) and specificity (95-99%) of EHR-based type 2 diabetes phenotypes (compared with the gold standard ADA criteria via chart review) varied depending on the component criteria and timing of observations and measurements. DISCUSSION AND CONCLUSIONS: Researchers using EHR-based phenotype definitions should clearly specify the characteristics that comprise the definition, variations of ADA criteria, and how different phenotype definitions and components impact the patient populations retrieved and the intended application. Careful attention to phenotype definitions is critical if the promise of leveraging EHR data to improve individual and population health is to be fulfilled. Susan E. Spratt, Katherine Pereira, Bradi B. Granger, Bryan C. Batch, Matthew Phelan, Michael J. Pencina, Marie Lynn Miranda, L. Ebony Boulware, Joseph E. Lucas, Charlotte L. Nelson, Benjamin Neely, Benjamin Goldstein 0001, Pamela Barth, Rachel L. Richesson, Isaretta L. Riley, Leonor Corsino, Eugenia R. McPeek Hinz, Shelley A. Rusincovitch, Jennifer Green, Anna Beth Barton, Carly Kelley, Kristen Hyland, Monica Tang, Amanda Elliott, Ewa Ruel, Alexander Clark, Melanie Mabrey, Kay Lyn Morrissey, Jyothi Rao, Beatrice Hong, Marjorie Pierre-Louis, Katherine Kelly, Nicole E. Jelesoff |
J. Am. Medical Informatics Assoc. | 12 |
| 2015 | A Simulation Framework for Longitudinal Electronic Health Records Data
Matthew Phelan, Benjamin Goldstein 0001 |
AMIA | 2 |
| 2015 | Classifying individuals based on a densely captured sequence of vital signs: An example using repeated blood pressure measurements during hemodialysis treatment
Benjamin Goldstein 0001, Tara I. Chang, Wolfgang C. Winkelmayer |
J. Biomed. Informatics | 1 |
| 2013 | Changes during dialysis captured in electronic health records help predict near-term risk of sudden cardiac death
Benjamin Goldstein 0001, Wolfgang C. Winkelmayer, Themistocles L. Assimes |
AMIA | 1 |