EDBT 2026 Demo / reviewers in the wild / expert
Rebecca A. Hubbard
dblp:216/7694
· DBLP profile ↗
13ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0003-0879-0994ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesisabstractOBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge. Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprintsabstractOBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts. Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 21 |
| 2024 | Leveraging error-prone algorithm-derived phenotypes: Enhancing association studies for risk factors in EHR data
Jiayi Tong, Jessica Chubak, Thomas Lumley, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 5 |
| 2022 | Informative presence bias in analyses of electronic health records-derived data: a cautionary noteabstractOBJECTIVE: Electronic health record (EHR)-derived data are extensively used in health research. However, the pattern of patient interaction with the healthcare system can result in informative presence bias if those who have poorer health have more data recorded than healthier patients. We aimed to determine how informative presence affects bias across multiple scenarios informed by real-world healthcare utilization patterns. MATERIALS AND METHODS: We conducted an analysis of EHR data from a pediatric healthcare system as well as simulation studies to characterize conditions under which informative presence bias is likely to occur. This analysis extends prior work by examining a variety of scenarios for the relationship between a biomarker and a health event of interest and the healthcare visit process. RESULTS: Using biomarker values gathered at both informative and noninformative visits when estimating the effect of the biomarker on the event of interest resulted in minimal bias when the biomarker was relatively stable over time but produced substantial bias when the biomarker was more volatile. Adjusting analyses for the number of prior visits within a fixed look-back window was able to reduce but not eliminate this bias. DISCUSSION: These results suggest that bias may arise frequently in commonly encountered scenarios and may not be eliminated by adjusting for prior visit intensity. CONCLUSION: Depending on the context, the estimated effect from analyses using data from all visits available may diverge from the true effect. Sensitivity analyses using only visits likely to be informative or noninformative based on visit type may aid in the assessment of the magnitude of potential bias. Joanna Harton, Nandita A. Mitra, Rebecca A. Hubbard |
J. Am. Medical Informatics Assoc. | 3 |
| 2022 | SAT: a Surrogate-Assisted Two-wave case boosting sampling method, with application to EHR-based association studiesabstractOBJECTIVES: Electronic health records (EHRs) enable investigation of the association between phenotypes and risk factors. However, studies solely relying on potentially error-prone EHR-derived phenotypes (ie, surrogates) are subject to bias. Analyses of low prevalence phenotypes may also suffer from poor efficiency. Existing methods typically focus on one of these issues but seldom address both. This study aims to simultaneously address both issues by developing new sampling methods to select an optimal subsample to collect gold standard phenotypes for improving the accuracy of association estimation. MATERIALS AND METHODS: We develop a surrogate-assisted two-wave (SAT) sampling method, where a surrogate-guided sampling (SGS) procedure and a modified optimal subsampling procedure motivated from A-optimality criterion (OSMAC) are employed sequentially, to select a subsample for outcome validation through manual chart review subject to budget constraints. A model is then fitted based on the subsample with the true phenotypes. Simulation studies and an application to an EHR dataset of breast cancer survivors are conducted to demonstrate the effectiveness of SAT. RESULTS: We found that the subsample selected with the proposed method contains informative observations that effectively reduce the mean squared error of the resultant estimator of the association. CONCLUSIONS: The proposed approach can handle the problem brought by the rarity of cases and misclassification of the surrogate in phenotype-absent EHR-based association studies. With a well-behaved surrogate, SAT successfully boosts the case prevalence in the subsample and improves the efficiency of estimation. Jessica Chubak, Rebecca A. Hubbard, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Development and validation of a prediction model for actionable aspects of frailty in the text of clinicians' encounter notesabstractOBJECTIVE: Frailty is a prevalent risk factor for adverse outcomes among patients with chronic lung disease. However, identifying frail patients who may benefit from interventions is challenging using standard data sources. We therefore sought to identify phrases in clinical notes in the electronic health record (EHR) that describe actionable frailty syndromes. MATERIALS AND METHODS: We used an active learning strategy to select notes from the EHR and annotated each sentence for 4 actionable aspects of frailty: respiratory impairment, musculoskeletal problems, fall risk, and nutritional deficiencies. We compared the performance of regression, tree-based, and neural network models to predict the labels for each sentence. We evaluated performance with the scaled Brier score (SBS), where 1 is perfect and 0 is uninformative, and the positive predictive value (PPV). RESULTS: We manually annotated 155 952 sentences from 326 patients. Elastic net regression had the best performance across all 4 frailty aspects (SBS 0.52, 95% confidence interval [CI] 0.49-0.54) followed by random forests (SBS 0.49, 95% CI 0.47-0.51), and multi-task neural networks (SBS 0.39, 95% CI 0.37-0.42). For the elastic net model, the PPV for identifying the presence of respiratory impairment was 54.8% (95% CI 53.3%-56.6%) at a sensitivity of 80%. DISCUSSION: Classification models using EHR notes can effectively identify actionable aspects of frailty among patients living with chronic lung disease. Regression performed better than random forest and neural network models. CONCLUSIONS: NLP-based models offer promising support to population health management programs that seek to identify and refer community-dwelling patients with frailty for evidence-based interventions. Jacob A. Martin, Andrew Crane-Droesch, Folasade C. Lapite, Joseph C. Puhl, Tyler E. Kmiec, Jasmine A. Silvestri, Lyle H. Ungar, Bruce P. Kinosian, Blanca E. Himes, Rebecca A. Hubbard, Joshua M. Diamond, Vivek N. Ahya, Michael W. Sims, Scott D. Halpern, Gary E. Weissman |
J. Am. Medical Informatics Assoc. | 10 |
| 2021 | A cost-effective chart review sampling design to account for phenotyping error in electronic health records (EHR) dataabstractOBJECTIVES: Electronic health records (EHR) are commonly used for the identification of novel risk factors for disease, often referred to as an association study. A major challenge to EHR-based association studies is phenotyping error in EHR-derived outcomes. A manual chart review of phenotypes is necessary for unbiased evaluation of risk factor associations. However, this process is time-consuming and expensive. The objective of this paper is to develop an outcome-dependent sampling approach for designing manual chart review, where EHR-derived phenotypes can be used to guide the selection of charts to be reviewed in order to maximize statistical efficiency in the subsequent estimation of risk factor associations. MATERIALS AND METHODS: After applying outcome-dependent sampling, an augmented estimator can be constructed by optimally combining the chart-reviewed phenotypes from the selected patients with the error-prone EHR-derived phenotype. We conducted simulation studies to evaluate the proposed method and applied our method to data on colon cancer recurrence in a cohort of patients treated for a primary colon cancer in the Kaiser Permanente Washington (KPW) healthcare system. RESULTS: Simulations verify the coverage probability of the proposed method and show that, when disease prevalence is less than 30%, the proposed method has smaller variance than an existing method where the validation set for chart review is uniformly sampled. In addition, from design perspective, the proposed method is able to achieve the same statistical power with 50% fewer charts to be validated than the uniform sampling method, thus, leading to a substantial efficiency gain in chart review. These findings were also confirmed by the application of the competing methods to the KPW colon cancer data. DISCUSSION: Our simulation studies and analysis of data from KPW demonstrate that, compared to an existing uniform sampling method, the proposed outcome-dependent method can lead to a more efficient chart review sampling design and unbiased association estimates with higher statistical efficiency. CONCLUSION: The proposed method not only optimally combines phenotypes from chart review with EHR-derived phenotypes but also suggests an efficient design for conducting chart review, with the goal of improving the efficiency of estimated risk factor associations using EHR data. Ziyan Yin, Jiayi Tong, Yong Chen 0016, Rebecca A. Hubbard, Cheng Yong Tang |
J. Am. Medical Informatics Assoc. | 4 |
| 2021 | Studying pediatric health outcomes with electronic health records using Bayesian clustering and trajectory analysisabstractUse of routinely collected data from electronic health records (EHR) can expedite longitudinal studies that investigate childhood exposures and rare pediatric health outcomes. For instance, characteristics of the body mass index (BMI) trajectory early in life may be associated with subsequent development of type 2 diabetes. Past studies investigating these relationships have used longitudinal cohort data collected over the course of many years to investigate the connection between BMI trajectory and subsequent development of diabetes. In contrast, EHR data from routine clinical care can provide longitudinal information on early-life BMI trajectories as well as subsequent health outcomes without requiring any additional data collection. In this study, we introduce a Bayesian joint phenotyping and BMI trajectory model to address data quality challenges in an EHR-based study of early-life BMI and type 2 diabetes in adolescence. We compared this joint modeling approach to traditional approaches using a computable phenotype for type 2 diabetes or separately estimated BMI trajectories and type 2 diabetes phenotypes. In a sample of 49,062 children derived from the PEDSnet consortium of pediatric healthcare systems, a median 8 (interquartile range [IQR] 5-13) BMI measurements were available to characterize the early-life BMI trajectory. The joint modeling and computable phenotype approaches found that age at adiposity rebound between 5 and 9 years was associated with higher odds of type 2 diabetes in adolescence compared to age at adiposity rebound between 2 and 5 years (joint model odds ratio [OR] = 1.77; computable phenotype OR = 1.88) and that BMI in excess of 140% of the 95th percentile for age and sex at age 9 years was associated with higher odds of type 2 diabetes in adolescence relative to children with BMI from 100 to 120% of the 95th percentile (joint model OR = 6.22; computable phenotype OR = 13.25). Estimates from the separate phenotyping and trajectory model were substantially attenuated towards the null. These results demonstrate that EHR data coupled with modern methodologic approaches can improve efficiency and timeliness of studies of childhood exposures and rare health outcomes. Rebecca A. Hubbard, Robert Siegel, Yong Chen 0016, Ihuoma Eneli |
J. Biomed. Informatics | 1 |
| 2020 | Impact of Individual versus Geographic-Area Measures of Socioeconomic Status on Health Associations Observed in the Behavioral Risk Factor Surveillance System
Lena Leszinsky, Sherrie Xie, Avantika Diwadkar, Rebecca Greenblatt, Rebecca A. Hubbard, Blanca E. Himes |
AMIA | 5 |
| 2020 | Identifying Actionable Aspects of Frailty in the Text of Encounter Notes
Jacob A. Martin, Andrew Crane-Droesch, Folasade C. Lapite, Joseph C. Puhl, Jasmine A. Silvestri, Bruce P. Kinosian, Blanca E. Himes, Rebecca A. Hubbard, Vivek N. Ahya, Michael W. Sims, Joshua M. Diamond, Joseph Adler, Elizabeth Steele, Emily Ott, Lyle H. Ungar, Scott D. Halpern, Gary E. Weissman |
AMIA | 8 |
| 2020 | An augmented estimation procedure for EHR-based association studies accounting for differential misclassificationabstractOBJECTIVES: The ability to identify novel risk factors for health outcomes is a key strength of electronic health record (EHR)-based research. However, the validity of such studies is limited by error in EHR-derived phenotypes. The objective of this study was to develop a novel procedure for reducing bias in estimated associations between risk factors and phenotypes in EHR data. MATERIALS AND METHODS: The proposed method combines the strengths of a gold-standard phenotype obtained through manual chart review for a small validation set of patients and an automatically-derived phenotype that is available for all patients but is potentially error-prone (hereafter referred to as the algorithm-derived phenotype). An augmented estimator of associations is obtained by optimally combining these 2 phenotypes. We conducted simulation studies to evaluate the performance of the augmented estimator and conducted an analysis of risk factors for second breast cancer events using data on a cohort from Kaiser Permanente Washington. RESULTS: The proposed method was shown to reduce bias relative to an estimator using only the algorithm-derived phenotype and reduce variance compared to an estimator using only the validation data. DISCUSSION: Our simulation studies and real data application demonstrate that, compared to the estimator using validation data only, the augmented estimator has lower variance (ie, higher statistical efficiency). Compared to the estimator using error-prone EHR-derived phenotypes, the augmented estimator has smaller bias. CONCLUSIONS: The proposed estimator can effectively combine an error-prone phenotype with gold-standard data from a limited chart review in order to improve analyses of risk factors using EHR data. Jiayi Tong, Jing Huang 0021, Jessica Chubak, Jason H. Moore, Rebecca A. Hubbard, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Analysis of Spatial Trends in Smoking Status Among Patients with Obstructive Airway Diseases Highlight Potential for Targeted Interventions
Sherrie Xie, Rebecca A. Hubbard, Blanca E. Himes |
AMIA | 2 |
| 2018 | PIE: A prior knowledge guided integrated likelihood estimation method for bias reduction in association studies using electronic health records dataabstractOBJECTIVES: This study proposes a novel Prior knowledge guided Integrated likelihood Estimation (PIE) method to correct bias in estimations of associations due to misclassification of electronic health record (EHR)-derived binary phenotypes, and evaluates the performance of the proposed method by comparing it to 2 methods in common practice. METHODS: We conducted simulation studies and data analysis of real EHR-derived data on diabetes from Kaiser Permanente Washington to compare the estimation bias of associations using the proposed method, the method ignoring phenotyping errors, the maximum likelihood method with misspecified sensitivity and specificity, and the maximum likelihood method with correctly specified sensitivity and specificity (gold standard). The proposed method effectively leverages available information on phenotyping accuracy to construct a prior distribution for sensitivity and specificity, and incorporates this prior information through the integrated likelihood for bias reduction. RESULTS: Our simulation studies and real data application demonstrated that the proposed method effectively reduces the estimation bias compared to the 2 current methods. It performed almost as well as the gold standard method when the prior had highest density around true sensitivity and specificity. The analysis of EHR data from Kaiser Permanente Washington showed that the estimated associations from PIE were very close to the estimates from the gold standard method and reduced bias by 60%-100% compared to the 2 commonly used methods in current practice for EHR data. CONCLUSIONS: This study demonstrates that the proposed method can effectively reduce estimation bias caused by imperfect phenotyping in EHR-derived data by incorporating prior information through integrated likelihood. Jing Huang 0021, Rui Duan 0004, Rebecca A. Hubbard, Yonghui Wu 0001, Jason H. Moore, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 3 |