VLDB 2026 Research / reviewers in the wild / expert
Elizabeth Shenkman
dblp:190/3489 · also Elizabeth A. Shenkman
· DBLP profile ↗
14ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0003-4903-1804ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 87 |
| 2024 | Evaluating site-of-care-related racial disparities in kidney graft failure using a novel federated learning frameworkabstractOBJECTIVES: Racial disparities in kidney transplant access and posttransplant outcomes exist between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in the United States, with the site of care being a key contributor. Using multi-site data to examine the effect of site of care on racial disparities, the key challenge is the dilemma in sharing patient-level data due to regulations for protecting patients' privacy. MATERIALS AND METHODS: We developed a federated learning framework, named dGEM-disparity (decentralized algorithm for Generalized linear mixed Effect Model for disparity quantification). Consisting of 2 modules, dGEM-disparity first provides accurately estimated common effects and calibrated hospital-specific effects by requiring only aggregated data from each center and then adopts a counterfactual modeling approach to assess whether the graft failure rates differ if NHB patients had been admitted at transplant centers in the same distribution as NHW patients were admitted. RESULTS: Utilizing United States Renal Data System data from 39 043 adult patients across 73 transplant centers over 10 years, we found that if NHB patients had followed the distribution of NHW patients in admissions, there would be 38 fewer deaths or graft failures per 10 000 NHB patients (95% CI, 35-40) within 1 year of receiving a kidney transplant on average. DISCUSSION: The proposed framework facilitates efficient collaborations in clinical research networks. Additionally, the framework, by using counterfactual modeling to calculate the event rate, allows us to investigate contributions to racial disparities that may occur at the level of site of care. CONCLUSIONS: Our framework is broadly applicable to other decentralized datasets and disparities research related to differential access to care. Ultimately, our proposed framework will advance equity in human health by identifying and addressing hospital-level racial disparities. Jiayi Tong, Yishan Shen, Alice Xu, Xing He 0003, Chongliang Luo, Mackenzie J. Edmondson, Dazheng Zhang, Chao Yan 0004, Ruowang Li, Lianne Siegel, Lichao Sun 0001, Elizabeth Shenkman, Sally C. Morton, Bradley A. Malin, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 13 |
| 2023 | The role of health system penetration rate in estimating the prevalence of type 1 diabetes in children and adolescents using electronic health recordsabstractOBJECTIVE: Having sufficient population coverage from the electronic health records (EHRs)-connected health system is essential for building a comprehensive EHR-based diabetes surveillance system. This study aimed to establish an EHR-based type 1 diabetes (T1D) surveillance system for children and adolescents across racial and ethnic groups by identifying the minimum population coverage from EHR-connected health systems to accurately estimate T1D prevalence. MATERIALS AND METHODS: We conducted a retrospective, cross-sectional analysis involving children and adolescents <20 years old identified from the OneFlorida+ Clinical Research Network (2018-2020). T1D cases were identified using a previously validated computable phenotyping algorithm. The T1D prevalence for each ZIP Code Tabulation Area (ZCTA, 5 digits), defined as the number of T1D cases divided by the total number of residents in the corresponding ZCTA, was calculated. Population coverage for each ZCTA was measured using observed health system penetration rates (HSPR), which was calculated as the ratio of residents in the corresponding ZTCA and captured by OneFlorida+ to the overall population in the same ZCTA reported by the Census. We used a recursive partitioning algorithm to identify the minimum required observed HSPR to estimate T1D prevalence and compare our estimate with the reported T1D prevalence from the SEARCH study. RESULTS: Observed HSPRs of 55%, 55%, and 60% were identified as the minimum thresholds for the non-Hispanic White, non-Hispanic Black, and Hispanic populations. The estimated T1D prevalence for non-Hispanic White and non-Hispanic Black were 2.87 and 2.29 per 1000 youth, which are comparable to the reference study's estimation. The estimated prevalence of T1D for Hispanics (2.76 per 1000 youth) was higher than the reference study's estimation (1.48-1.64 per 1000 youth). The standardized T1D prevalence in the overall Florida population was 2.81 per 1000 youth in 2019. CONCLUSION: Our study provides a method to estimate T1D prevalence in children and adolescents using EHRs and reports the estimated HSPRs and prevalence of T1D for different race and ethnicity groups to facilitate EHR-based diabetes surveillance. Piaopiao Li, Tianchen Lyu, Khalid Alkhuzam, Eliot Spector, William T. Donahoo, Sarah Bost, Yonghui Wu 0001, William R. Hogan, Mattia Prosperi, Desmond A. Schatz, Mark A. Atkinson, Michael J. Haller, Elizabeth Shenkman, Yi Guo 0005, Jiang Bian 0001 |
J. Am. Medical Informatics Assoc. | 13 |
| 2022 | The OneFlorida Data Trust: a centralized, translational research data infrastructure of statewide scopeabstractThe OneFlorida Data Trust is a centralized research patient data repository created and managed by the OneFlorida Clinical Research Consortium ("OneFlorida"). It comprises structured electronic health record (EHR), administrative claims, tumor registry, death, and other data on 17.2 million individuals who received healthcare in Florida between January 2012 and the present. Ten healthcare systems in Miami, Orlando, Tampa, Jacksonville, Tallahassee, Gainesville, and rural areas of Florida contribute EHR data, covering the major metropolitan regions in Florida. Deduplication of patients is accomplished via privacy-preserving entity resolution (precision 0.97-0.99, recall 0.75), thereby linking patients' EHR, claims, and death data. Another unique feature is the establishment of mother-baby relationships via Florida vital statistics data. Research usage has been significant, including major studies launched in the National Patient-Centered Clinical Research Network ("PCORnet"), where OneFlorida is 1 of 9 clinical research networks. The Data Trust's robust, centralized, statewide data are a valuable and relatively unique research resource. William R. Hogan, Elizabeth Shenkman, Temple Robinson, Olveen Carasquillo, Patricia S. Robinson, Rebecca Z. Essner, Jiang Bian 0001, Gigi Lipori, Christopher A. Harle, Tanja Magoc, Lizabeth Manini, Tona Mendoza, Sonya White, Alex Loiacono, Jackie Hall, Dave Nelson |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Developing and Validating a Computable Phenotype for the Identification of Transgender and Gender Nonconforming Individuals and Subgroups
Yi Guo 0005, Xing He 0003, Tianchen Lyu, Hansi Zhang, Yonghui Wu 0001, Xi Yang 0015, Zhaoyi Chen, Merry J. Markham, François Modave, Mengjun Xie, William R. Hogan, Christopher A. Harle, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 13 |
| 2020 | Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world dataabstractOBJECTIVE: To synthesize data quality (DQ) dimensions and assessment methods of real-world data, especially electronic health records, through a systematic scoping review and to assess the practice of DQ assessment in the national Patient-centered Clinical Research Network (PCORnet). MATERIALS AND METHODS: We started with 3 widely cited DQ literature-2 reviews from Chan et al (2010) and Weiskopf et al (2013a) and 1 DQ framework from Kahn et al (2016)-and expanded our review systematically to cover relevant articles published up to February 2020. We extracted DQ dimensions and assessment methods from these studies, mapped their relationships, and organized a synthesized summarization of existing DQ dimensions and assessment methods. We reviewed the data checks employed by the PCORnet and mapped them to the synthesized DQ dimensions and methods. RESULTS: We analyzed a total of 3 reviews, 20 DQ frameworks, and 226 DQ studies and extracted 14 DQ dimensions and 10 assessment methods. We found that completeness, concordance, and correctness/accuracy were commonly assessed. Element presence, validity check, and conformance were commonly used DQ assessment methods and were the main focuses of the PCORnet data checks. DISCUSSION: Definitions of DQ dimensions and methods were not consistent in the literature, and the DQ assessment practice was not evenly distributed (eg, usability and ease-of-use were rarely discussed). Challenges in DQ assessments, given the complex and heterogeneous nature of real-world data, exist. CONCLUSION: The practice of DQ assessment is still limited in scope. Future work is warranted to generate understandable, executable, and reusable DQ measures. Jiang Bian 0001, Tianchen Lyu, Alexander T. Loiacono, Tonatiuh Mendoza Viramontes, Gloria P. Lipori, Yi Guo 0005, Yonghui Wu 0001, Mattia Prosperi, Thomas J. George, Christopher A. Harle, Elizabeth Shenkman, William R. Hogan |
J. Am. Medical Informatics Assoc. | 11 |
| 2019 | If You Build It, They Will Come: The National Patient-Centered Clinical Research Network (PCORnet): From Conception to Execution
Thomas Carton, Maryan Zirkle, Elizabeth Shenkman, Abel N. Kho, Adrian Hernandez |
AMIA | 3 |
| 2018 | Clustering Inter-Arrival Time of Health Care Encounters for High UtilizersabstractPatients with a large number of health care encounters receive great attention in health care research because their expensive and problematic health care utilization has important implications for the US health care system. The large volume of emergency department (ED) visits and inpatient hospital stays through time provides a unique opportunity to apply data-driven methods for identifying temporal signals associated with these patients. The micro variations in the inter-arrival time of these encounters within this population are not well-studied. Computational approaches for distinguishing various temporal visiting patterns leads to the problem of efficiently clustering asynchronous time series. Thus, we propose a Wasserstein distance based spectral clustering for this problem. Asynchronous time series are first represented as histograms of inter-arrival time and their pairwise similarities are computed under Wasserstein distance. Spectral clustering operates on the similarity matrix as input thereby avoiding the computational bottleneck of Wasserstein barycenters. The effectiveness of this method is demonstrated by synthetic data and application to a large real world health insurance encounters dataset, identifying potential associations between temporal visiting patterns and health factors from a population of frequent ED and inpatient hospital users. Chengliang Yang, Chris Delcher, Elizabeth Shenkman, Sanjay Ranka |
HealthCom | 3 |
| 2017 | Comparing and Contrasting A Priori and A Posteriori Generalizability Assessment of Clinical Trials on Type 2 Diabetes Mellitus
Zhe He 0001, Arturo Gonzalez-Izquierdo, Spiros C. Denaxas, Andrei Sura, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
AMIA | 7 |
| 2017 | Implementing a Hash-based Privacy-Preserving Entity Resolution Tool in the OneFlorida Clinical Data Research Network
Jiang Bian 0001, Andrei Sura, Gloria P. Lipori, Yi Guo 0005, François Modave, Zhe He 0001, Elizabeth Shenkman, William R. Hogan |
AMIA | 7 |
| 2017 | Identifying High Health Care Utilizers Using Post-Regression Residual Analysis of Health Expenditures from a State Medicaid Program
Chengliang Yang, Chris Delcher, Elizabeth Shenkman, Sanjay Ranka |
AMIA | 3 |
| 2017 | Data integration through ontology-based data access to support integrative data analysis: A case study of cancer survivalabstractTo improve cancer survival rates and prognosis, one of the first steps is to improve our understanding of contributory factors associated with cancer survival. Prior research has suggested that cancer survival is influenced by multiple factors from multiple levels. Most of existing analyses of cancer survival used data from a single source. Nevertheless, there are key challenges in integrating variables from different sources. Data integration is a daunting task because data from different sources can be heterogeneous in syntax, schema, and particularly semantics. Thus, we propose to adopt a semantic data integration approach that generates a universal conceptual representation of "information" including data and their relationships. This paper describes a case study of semantic data integration linking three data sets that cover both individual and contextual level factors for the purpose of assessing the association of the predictors of interest with cancer survival using cox proportional hazard models. Hansi Zhang, Yi Guo 0005, Qian Li 0034, Thomas J. George, Elizabeth Shenkman, Jiang Bian 0001 |
BIBM | 5 |
| 2017 | Towards a privacy preserving cohort discovery framework for clinical research networks
Bradley A. Malin, François Modave, Yi Guo 0005, William R. Hogan, Elizabeth Shenkman, Jiang Bian 0001 |
J. Biomed. Informatics | 6 |
| 2016 | Predicting 30-day all-cause readmissions from hospital inpatient discharge dataabstractInpatient hospital readmissions for potentially avoidable conditions are problematic and costly. In this paper, we build machine learning models using variables widely available in health claims data to predict patients' 30-day readmission risks at the time of discharge. These models show high predictive power on a U.S. nationwide readmission database. They are also capable of providing interpretable risk factors globally at the population level and locally associated with each single discharge. In addition, we propose a model-agnostic approach to provide confidence for each prediction. Altogether, using models with high predictive power, interpretable risk factors and prediction confidence may enable health care systems to accurately target high-risk patients and prevent recurrent readmissions by accurately anticipating the probability of readmission at the point of care. Chengliang Yang, Chris Delcher, Elizabeth Shenkman, Sanjay Ranka |
HealthCom | 3 |