EDBT 2026 Demo / reviewers in the wild / expert
Jiayi Tong
dblp:234/3273
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference chartsabstractMOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS. Fengling Hu, Jiayi Tong, Margaret Gardner, Lifespan Brain Chart Consortium, Andrew A. Chen, Richard A. I. Bethlehem, Jakob Seidlitz, Hongzhe Li, Aaron Alexander-Bloch, Yong Chen 0016, Russell T. Shinohara |
Bioinform. | 2 |
| 2026 | A lossless one-shot distributed algorithm for addressing heterogeneity in multi-site generalized linear modelsabstractOBJECTIVE: We propose Heterogeneity-aware Collaborative One-shot Lossless Algorithm for Generalized Linear Model (COLA-GLM-H), a novel one-shot lossless distributed algorithm that enables the integration of heterogeneous multi-institutional data while relying solely on instituion-level summary information rather than patient-level data. MATERIALS AND METHODS: Generalized Linear Models (GLMs) are widely used in medical research for analyzing diverse outcome types. In multi-institution settings, we demonstrated that the global likelihood can be reconstructed using only institution-level summary statistics, enabling lossless estimation without accessing individual records. We validated COLA-GLM-H in two real-world studies: (1) an emulated U.S. pediatric centralized network (719,383 patients) evaluating long-term cardiovascular risks following COVID-19, and (2) an internationally decentralized network of 120,429 hospitalized patients from seven databases across three countries assessing risk factors for COVID-19 mortality. RESULTS: In the centralized network, COLA-GLM-H produced estimates identical to those from pooled analyses. In the decentralized setting, the algorithm effectively integrated heterogeneous data across multiple clinical institutions using a single communication round. CONCLUSIONS: COLA-GLM-H provides a lossless, communication-efficient, and computation-efficient solution for multi-institutional research using only institution-level summary data. It accounts for between-institution heterogeneity and supports all outcome types within the exponential family, enabling secure, scalable, and accurate analysis in collaborative clinical research. Bingyu Zhang, Jenna Reps, Jiayi Tong, Dazheng Zhang, Juan Manuel Ramírez-Anguita, Jiang Bian 0001, Milou T. Brand, Thomas Falconer, Miguel A. Mayer, Ross D. Williams, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 5 |
| 2026 | A multidimensional hierarchical framework for sources of bias in real-world healthcare evidence: a scoping review
Christelle Xiong, Derek Baughman, Chen Dun, Jiayi Tong, Harold P. Lehmann, Paul G. Nagy |
J. Biomed. Informatics | 5 |
| 2025 | Incorporating preprints in systematic reviews: a preliminary study of a novel method for rapid evidence synthesisabstractOBJECTIVES: By October 1, 2024, over 450,000 COVID-19 manuscripts were published, with 10% posted as unreviewed preprints. While they accelerate knowledge sharing, their inconsistent quality complicates systematic studies. MATERIALS AND METHODS: We propose a 2-stage method to include preprints in meta-analyses. In Stage A, preprints are integrated through restriction or imputation and weighted by a confidence score reflecting their publication likelihood. In Stage B, we assess and adjust for potential publication or reporting biases. RESULTS: This preliminary study employed a 2-stage procedure validated with 2 COVID-19 treatment case studies. For hydroxychloroquine, the relative risk (RR) was 1.06 [95% CI: 0.62, 1.80], suggesting no mortality benefit over placebo. For corticosteroids, the RR was 0.88 [95% CI: 0.62, 1.27], which, while not statistically significant, aligns with evidence supporting a mortality benefit. DISCUSSION: Our research aims to bridge a significant methodological gap by providing a solution for timely evidence synthesis, particularly in the face of the overwhelming number of publications surrounding COVID-19. CONCLUSION: This preliminary study presents a method to efficiently synthesize COVID-19 research, including non-peer-reviewed preprints, to support clinical and policy decisions amidst the information surge. Jiayi Tong, Yifei Sun 0007, Rebecca A. Hubbard, M. Elle Saine, Hua Xu 0001, Xu Zuo, Chunhua Weng, Christopher H. Schmid, Stephen E. Kimmel, Craig A. Umscheid, Adam Cuker, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 1 |
| 2025 | A communication-efficient federated learning algorithm to assess racial disparities in post-transplantation survival timeabstractOBJECTIVE: Patients of different race have different outcomes following renal transplantation. Patients of different race also undergo renal transplantation at different hospitals. We used a novel decentralized multisite approach to quantitatively assess the effect of site of care on racial disparities between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in post-transplantation survival times. MATERIALS AND METHODS: In this study, we develop a communication-efficient federated learning algorithm to assess site-of-care associated racial disparities based on decentralized time-to-event data, called Communication-Efficient Distributed Analysis for Racial Disparity in Time-to-event Data (CEDAR-t2e). The algorithm includes 2 modules. Module I is to estimate the site-specific proportional hazards model for time-to-event outcomes in a distributed manner, in which the Poissonization is used to simplify the estimation procedure. Based on the estimated results from Module I, Module II calculates how long the kidney failure time of NHB patients would be extended had they been admitted to transplant centers in the same distribution as NHW patients were admitted. RESULTS: With application to United States Renal Data System data covering 39 043 patients across 73 transplant centers, we found no evidence suggesting the presence of site-of-care associated racial disparities in post-transplantation survival times. In particular, restricting to one year after transplantation, the counterfactual graft failure time would have been extended by only 0.61 days on average if NHB had the same admission distribution to transplant centers as NHW patients. DISCUSSION: The proposed approach offers a quantitative measure to evaluate site-of-care associated racial disparities. CONCLUSION: Our approach has the potential to be extended to investigate site-of-care related disparities in other time-to-event outcomes, thus promoting health equity and improving patient health in various fields. Dazheng Zhang, Jiayi Tong, Xing He 0003, Liang Li 0026, Lichao Sun 0001, Ashutosh M. Shukla, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Enabling inclusive systematic reviews: incorporating preprint articles with large language model-driven evaluationsabstractOBJECTIVES: Systematic reviews in comparative effectiveness research require timely evidence synthesis. With the rapid advancement of medical research, preprint articles play an increasingly important role in accelerating knowledge dissemination. However, as preprint articles are not peer-reviewed before publication, their quality varies significantly, posing challenges for evidence inclusion in systematic reviews. MATERIALS AND METHODS: We developed AutoConfidenceScore (automated confidence score assessment), an advanced framework for predicting preprint publication, which reduces reliance on manual curation and expands the range of predictors, including three key advancements: (1) automated data extraction using natural language processing techniques, (2) semantic embeddings of titles and abstracts, and (3) large language model (LLM)-driven evaluation scores. Additionally, we employed two prediction models: a random forest classifier for binary outcome and a survival cure model that predicts both binary outcome and publication risk over time. RESULTS: The random forest classifier achieved an area under the receiver operating characteristic curve (AUROC) of 0.747 using all features. The survival cure model achieved an AUROC of 0.731 for binary outcome prediction and a concordance index of 0.667 for time-to-publication risk. DISCUSSION: Our study advances the framework for preprint publication prediction through automated data extraction and multiple feature integration. By combining semantic embeddings with LLM-driven evaluations, AutoConfidenceScore significantly enhances predictive performance while reducing manual annotation burden. CONCLUSION: AutoConfidenceScore has the potential to facilitate incorporation of preprint articles during the appraisal phase of systematic reviews, supporting researchers in more effective utilization of preprint resources. Rui Yang 0016, Jiayi Tong, Nan Liu 0003, Christopher J. Lindsell, Michael J. Pencina, Yong Chen 0016, Chuan Hong |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | Leveraging undecided cases in chart-reviewed phenotypes to enhance EHR-based association studies
Xinyao Jian, Dazheng Zhang, Zehao Yu 0001, Hua Xu 0001, Jiang Bian 0001, Yonghui Wu 0001, Jiayi Tong, Yong Chen 0016 |
J. Biomed. Informatics | 7 |
| 2025 | Evaluating the Bias, type I error and statistical power of the prior Knowledge-Guided integrated likelihood estimation (PIE) for bias reduction in EHR based association studiesabstract• Question: How does PIE perform in various types of real-world scenarios, in terms of estimation and hypothesis testing? • Findings: Under non-differential misclassification, PIE had a smaller bias in estimated associations compared to the naïve method, but it had similar type I error and power. • The bias reduction of PIE was superior when the prior distribution of sensitivity and specificity of the phenotyping algorithm is more accurate (i.e., close to the true operating characteristics of the phenotyping algorithm). The impact of prior is relatively small when the outcome has low prevalence and is larger when the outcome is common. • PIE can effectively reduce the bias due to phenotyping error under a wide spectrum of real-world settings. However, its main advantage is in the reduction of bias in estimation but not in hypothesis testing. Binary outcomes in electronic health records (EHR) derived using automated phenotype algorithms may suffer from phenotyping error, resulting in bias in association estimation. Huang et al. [1] proposed the Prior Knowledge-Guided Integrated Likelihood Estimation (PIE) method to mitigate the estimation bias, however, their investigation focused on point estimation without statistical inference, and the evaluation of PIE therein using simulation was a proof-of-concept with only a limited scope of scenarios. This study aims to comprehensively assess PIE’s performance including (1) how well PIE performs under a wide spectrum of operating characteristics of phenotyping algorithms under real-world scenarios (e. g., low prevalence, low sensitivity, high specificity); (2) beyond point estimation, how much variation of the PIE estimator was introduced by the prior distribution; and (3) from a hypothesis testing point of view, if PIE improves type I error and statistical power relative to the naïve method (i.e., ignoring the phenotyping error). Synthetic data and use-case analysis were utilized to evaluate PIE. The synthetic data were generated under diverse outcome prevalence, phenotyping algorithm sensitivity, and association effect sizes. Simulation studies compared PIE under different prior distributions with the naïve method, assessing bias, variance, type I error, and power. Use-case analysis compared the performance of PIE and the naïve method in estimating the association of multiple predictors with COVID-19 infection. PIE exhibited reduced bias compared to the naïve method across varied simulation settings, with comparable type I error and power. As the effect size became larger, the bias reduced by PIE was larger. PIE has superior performance when prior distributions aligned closely with true phenotyping algorithm characteristics. Impact of prior quality was minor for low-prevalence outcomes but large for common outcomes. In use-case analysis, PIE maintains a relatively accurate estimation across different scenarios, particularly outperforming the naïve approach under large effect sizes. PIE effectively mitigates estimation bias in a wide spectrum of real-world settings, particularly with accurate prior information. Its main benefit lies in bias reduction rather than hypothesis testing. The impact of the prior is small for low-prevalence outcomes. Naimin Jing, Jiayi Tong, James Weaver, Patrick B. Ryan, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 3 |
| 2025 | DisC2o-HD: Distributed causal inference with covariates shift for analyzing real-world high-dimensional dataabstractHigh-dimensional healthcare data, such as electronic health records (EHR) data and claims data, present two primary challenges due to the large number of variables and the need to consolidate data from multiple clinical sites. The third key challenge is the potential existence of heterogeneity in terms of covariate shift. In this paper, we propose a distributed learning algorithm accounting for covariate shift to estimate the average treatment effect (ATE) for high-dimensional data, named DisC2o-HD. Leveraging the surrogate likelihood method, our method calibrates the estimates of the propensity score and outcome models to approximately attain the desired covariate balancing property, while accounting for the covariate shift across multiple clinical sites. We show that our distributed covariate balancing propensity score estimator can approximate the pooled estimator, which is obtained by pooling the data from multiple sites together. The proposed estimator remains consistent if either the propensity score model or the outcome regression model is correctly specified. The semiparametric efficiency bound is achieved when both the propensity score and the outcome models are correctly specified. We conduct simulation studies to demonstrate the performance of the proposed algorithm; additionally, we conduct an empirical study to present the readiness of implementation and validity. Jiayi Tong, George Hripcsak, Yang Ning, Yong Chen 0016 |
J. Mach. Learn. Res. | 1 |
| 2024 | Confidence score: a data-driven measure for inclusive systematic reviews considering unpublished preprintsabstractOBJECTIVES: COVID-19, since its emergence in December 2019, has globally impacted research. Over 360 000 COVID-19-related manuscripts have been published on PubMed and preprint servers like medRxiv and bioRxiv, with preprints comprising about 15% of all manuscripts. Yet, the role and impact of preprints on COVID-19 research and evidence synthesis remain uncertain. MATERIALS AND METHODS: We propose a novel data-driven method for assigning weights to individual preprints in systematic reviews and meta-analyses. This weight termed the "confidence score" is obtained using the survival cure model, also known as the survival mixture model, which takes into account the time elapsed between posting and publication of a preprint, as well as metadata such as the number of first 2-week citations, sample size, and study type. RESULTS: Using 146 preprints on COVID-19 therapeutics posted from the beginning of the pandemic through April 30, 2021, we validated the confidence scores, showing an area under the curve of 0.95 (95% CI, 0.92-0.98). Through a use case on the effectiveness of hydroxychloroquine, we demonstrated how these scores can be incorporated practically into meta-analyses to properly weigh preprints. DISCUSSION: It is important to note that our method does not aim to replace existing measures of study quality but rather serves as a supplementary measure that overcomes some limitations of current approaches. CONCLUSION: Our proposed confidence score has the potential to improve systematic reviews of evidence related to COVID-19 and other clinical conditions by providing a data-driven approach to including unpublished manuscripts. Jiayi Tong, Chongliang Luo, Yifei Sun 0007, Rui Duan 0004, M. Elle Saine, Yifan Peng 0002, Anchita Batra, Anni Pan, Olivia Wang, Ruowang Li, Arielle Marks-Anglin, Xu Zuo, Yulun Liu 0004, Jiang Bian 0001, Stephen E. Kimmel, Keith Hamilton, Adam Cuker, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Evaluating site-of-care-related racial disparities in kidney graft failure using a novel federated learning frameworkabstractOBJECTIVES: Racial disparities in kidney transplant access and posttransplant outcomes exist between non-Hispanic Black (NHB) and non-Hispanic White (NHW) patients in the United States, with the site of care being a key contributor. Using multi-site data to examine the effect of site of care on racial disparities, the key challenge is the dilemma in sharing patient-level data due to regulations for protecting patients' privacy. MATERIALS AND METHODS: We developed a federated learning framework, named dGEM-disparity (decentralized algorithm for Generalized linear mixed Effect Model for disparity quantification). Consisting of 2 modules, dGEM-disparity first provides accurately estimated common effects and calibrated hospital-specific effects by requiring only aggregated data from each center and then adopts a counterfactual modeling approach to assess whether the graft failure rates differ if NHB patients had been admitted at transplant centers in the same distribution as NHW patients were admitted. RESULTS: Utilizing United States Renal Data System data from 39 043 adult patients across 73 transplant centers over 10 years, we found that if NHB patients had followed the distribution of NHW patients in admissions, there would be 38 fewer deaths or graft failures per 10 000 NHB patients (95% CI, 35-40) within 1 year of receiving a kidney transplant on average. DISCUSSION: The proposed framework facilitates efficient collaborations in clinical research networks. Additionally, the framework, by using counterfactual modeling to calculate the event rate, allows us to investigate contributions to racial disparities that may occur at the level of site of care. CONCLUSIONS: Our framework is broadly applicable to other decentralized datasets and disparities research related to differential access to care. Ultimately, our proposed framework will advance equity in human health by identifying and addressing hospital-level racial disparities. Jiayi Tong, Yishan Shen, Alice Xu, Xing He 0003, Chongliang Luo, Mackenzie J. Edmondson, Dazheng Zhang, Chao Yan 0004, Ruowang Li, Lianne Siegel, Lichao Sun 0001, Elizabeth Shenkman, Sally C. Morton, Bradley A. Malin, Jiang Bian 0001, David A. Asch, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 1 |
| 2024 | Learning competing risks across multiple hospitals: one-shot distributed algorithmsabstractOBJECTIVES: To characterize the complex interplay between multiple clinical conditions in a time-to-event analysis framework using data from multiple hospitals, we developed two novel one-shot distributed algorithms for competing risk models (ODACoR). By applying our algorithms to the EHR data from eight national children's hospitals, we quantified the impacts of a wide range of risk factors on the risk of post-acute sequelae of SARS-COV-2 (PASC) among children and adolescents. MATERIALS AND METHODS: Our ODACoR algorithms are effectively executed due to their devised simplicity and communication efficiency. We evaluated our algorithms via extensive simulation studies as applications to quantification of the impacts of risk factors for PASC among children and adolescents using data from eight children's hospitals including the Children's Hospital of Philadelphia, Cincinnati Children's Hospital Medical Center, Children's Hospital of Colorado covering over 6.5 million pediatric patients. The accuracy of the estimation was assessed by comparing the results from our ODACoR algorithms with the estimators derived from the meta-analysis and the pooled data. RESULTS: The meta-analysis estimator showed a high relative bias (∼40%) when the clinical condition is relatively rare (∼0.5%), whereas ODACoR algorithms exhibited a substantially lower relative bias (∼0.2%). The estimated effects from our ODACoR algorithms were identical on par with the estimates from the pooled data, suggesting the high reliability of our federated learning algorithms. In contrast, the meta-analysis estimate failed to identify risk factors such as age, gender, chronic conditions history, and obesity, compared to the pooled data. DISCUSSION: Our proposed ODACoR algorithms are communication-efficient, highly accurate, and suitable to characterize the complex interplay between multiple clinical conditions. CONCLUSION: Our study demonstrates that our ODACoR algorithms are communication-efficient and can be widely applicable for analyzing multiple clinical conditions in a time-to-event analysis framework. Dazheng Zhang, Jiayi Tong, Naimin Jing, Chongliang Luo, Dimitri A. Christakis, Diana Güthe, Mady Hornig, Kelly J. Kelleher, Keith E. Morse, Colin M. Rogerson, Jasmin Divers, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 2 |
| 2024 | Leveraging error-prone algorithm-derived phenotypes: Enhancing association studies for risk factors in EHR data
Jiayi Tong, Jessica Chubak, Thomas Lumley, Rebecca A. Hubbard, Hua Xu 0001, Yong Chen 0016 |
J. Biomed. Informatics | 2 |
| 2024 | One-shot distributed algorithms for addressing heterogeneity in competing risks data across clinical sites
Dazheng Zhang, Jiayi Tong, Ronen Stein, Naimin Jing, Mary Regina Boland, Chongliang Luo, Robert N. Baldassano, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Biomed. Informatics | 2 |
| 2021 | A cost-effective chart review sampling design to account for phenotyping error in electronic health records (EHR) dataabstractOBJECTIVES: Electronic health records (EHR) are commonly used for the identification of novel risk factors for disease, often referred to as an association study. A major challenge to EHR-based association studies is phenotyping error in EHR-derived outcomes. A manual chart review of phenotypes is necessary for unbiased evaluation of risk factor associations. However, this process is time-consuming and expensive. The objective of this paper is to develop an outcome-dependent sampling approach for designing manual chart review, where EHR-derived phenotypes can be used to guide the selection of charts to be reviewed in order to maximize statistical efficiency in the subsequent estimation of risk factor associations. MATERIALS AND METHODS: After applying outcome-dependent sampling, an augmented estimator can be constructed by optimally combining the chart-reviewed phenotypes from the selected patients with the error-prone EHR-derived phenotype. We conducted simulation studies to evaluate the proposed method and applied our method to data on colon cancer recurrence in a cohort of patients treated for a primary colon cancer in the Kaiser Permanente Washington (KPW) healthcare system. RESULTS: Simulations verify the coverage probability of the proposed method and show that, when disease prevalence is less than 30%, the proposed method has smaller variance than an existing method where the validation set for chart review is uniformly sampled. In addition, from design perspective, the proposed method is able to achieve the same statistical power with 50% fewer charts to be validated than the uniform sampling method, thus, leading to a substantial efficiency gain in chart review. These findings were also confirmed by the application of the competing methods to the KPW colon cancer data. DISCUSSION: Our simulation studies and analysis of data from KPW demonstrate that, compared to an existing uniform sampling method, the proposed outcome-dependent method can lead to a more efficient chart review sampling design and unbiased association estimates with higher statistical efficiency. CONCLUSION: The proposed method not only optimally combines phenotypes from chart review with EHR-derived phenotypes but also suggests an efficient design for conducting chart review, with the goal of improving the efficiency of estimated risk factor associations using EHR data. Ziyan Yin, Jiayi Tong, Yong Chen 0016, Rebecca A. Hubbard, Cheng Yong Tang |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Leverage Real-World Longitudinal Data in Large Clinical Research Networks for Alzheimer's Disease and Related Dementia (ADRD)
Rui Duan 0004, Zhaoyi Chen, Jiayi Tong, Chongliang Luo, Tianchen Lyu, Cui Tao, Demetrius Maraganore, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 3 |
| 2020 | Identifying Clinical Risk Factors for Opioid Use Disorder using a Distributed Algorithm to Combine Real-World Data from a Large Clinical Data Research Network
Jiayi Tong, Zhaoyi Chen, Rui Duan 0004, Wei-Hsuan Lo-Ciganic, Tianchen Lyu, Cui Tao, Peter A. Merkel, Henry R. Kranzler, Jiang Bian 0001, Yong Chen 0016 |
AMIA | 1 |
| 2020 | Learning from local to global: An efficient distributed algorithm for modeling time-to-event dataabstractOBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner. Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | An augmented estimation procedure for EHR-based association studies accounting for differential misclassificationabstractOBJECTIVES: The ability to identify novel risk factors for health outcomes is a key strength of electronic health record (EHR)-based research. However, the validity of such studies is limited by error in EHR-derived phenotypes. The objective of this study was to develop a novel procedure for reducing bias in estimated associations between risk factors and phenotypes in EHR data. MATERIALS AND METHODS: The proposed method combines the strengths of a gold-standard phenotype obtained through manual chart review for a small validation set of patients and an automatically-derived phenotype that is available for all patients but is potentially error-prone (hereafter referred to as the algorithm-derived phenotype). An augmented estimator of associations is obtained by optimally combining these 2 phenotypes. We conducted simulation studies to evaluate the performance of the augmented estimator and conducted an analysis of risk factors for second breast cancer events using data on a cohort from Kaiser Permanente Washington. RESULTS: The proposed method was shown to reduce bias relative to an estimator using only the algorithm-derived phenotype and reduce variance compared to an estimator using only the validation data. DISCUSSION: Our simulation studies and real data application demonstrate that, compared to the estimator using validation data only, the augmented estimator has lower variance (ie, higher statistical efficiency). Compared to the estimator using error-prone EHR-derived phenotypes, the augmented estimator has smaller bias. CONCLUSIONS: The proposed estimator can effectively combine an error-prone phenotype with gold-standard data from a limited chart review in order to improve analyses of risk factors using EHR data. Jiayi Tong, Jing Huang 0021, Jessica Chubak, Jason H. Moore, Rebecca A. Hubbard, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Identification of Rare Adverse Events with Year-varying Reporting Rates for FLU4 Vaccine in VAERS
Jiayi Tong, Jing Huang 0021, Jingcheng Du, Cui Tao, Yong Chen 0016 |
AMIA | 1 |
| 2018 | Comparing adverse effects of Hepatitis C drugs using FAERS data
Jing Huang 0021, Xinyuan Zhang 0003, Jiayi Tong, Jingcheng Du, Rui Duan 0004, Liu Yang 0026, Jason H. Moore, Yong Chen 0016, Cui Tao |
BIBM | 3 |