VLDB 2026 Research / reviewers in the wild / expert
Christopher B. Forrest
dblp:91/9252
· DBLP profile ↗
13ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0003-1252-068XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysisabstractOBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases. Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor |
J. Am. Medical Informatics Assoc. | 24 |
| 2026 | A multifaceted approach to advancing data quality and fitness standards in multi-institutional networksabstractOBJECTIVE: To construct a data quality (DQ) system that incorporates combinations of methods to evaluate data characteristics and analytic fitness across research questions for multiple uses. MATERIALS AND METHODS: Drawing from experience of other data quality programs, network data extraction needs, and recurring study requirements, we developed 5 standards to guide development of a modular, multifaceted data quality system. These included annotation and documentation, ability to measure research readiness, reproducibility across networks, flexibility for the user, and interpretability to research and project teams. Implementation of checks based on these principles focused on reusability and interactive visualization of results. RESULTS: We identified 10 check types producing over 444 check applications and deployed them in 2 multi-institutional networks. Check types span structural conformance to a data model, utility for common research needs, and study-specific customization. All check types are customizable without dependencies between them. A dashboard visualizes results, permitting adjustments based on number of data sources, need for source masking, and the user's focus. All components can be applied as written to any data source using OMOP and are readily modified for other data models. DISCUSSION: We have extended previous work through our novel and multifaceted approach to data quality assessment, addressing needs in both network data improvement and research usage. We developed a capable and deployable system rather than tailoring to specific use cases. CONCLUSION: Our novel DQ assessment system provides essential components for future standardization and collaboration to improve fitness of clinical data for intended use. Hanieh Razzaghi, Kimberley Dickinson, Kaleigh Wieand, Samuel Boss, Hunter Weidlich, Yungui Huang, Keith E. Morse, Sujan Kumar Mutyala, Jyothi Priya Alekapatti Nandagopal, Karthik Viswanathan, Christopher B. Forrest, L. Charles Bailey |
J. Am. Medical Informatics Assoc. | 11 |
| 2024 | Learning competing risks across multiple hospitals: one-shot distributed algorithmsabstractOBJECTIVES: To characterize the complex interplay between multiple clinical conditions in a time-to-event analysis framework using data from multiple hospitals, we developed two novel one-shot distributed algorithms for competing risk models (ODACoR). By applying our algorithms to the EHR data from eight national children's hospitals, we quantified the impacts of a wide range of risk factors on the risk of post-acute sequelae of SARS-COV-2 (PASC) among children and adolescents. MATERIALS AND METHODS: Our ODACoR algorithms are effectively executed due to their devised simplicity and communication efficiency. We evaluated our algorithms via extensive simulation studies as applications to quantification of the impacts of risk factors for PASC among children and adolescents using data from eight children's hospitals including the Children's Hospital of Philadelphia, Cincinnati Children's Hospital Medical Center, Children's Hospital of Colorado covering over 6.5 million pediatric patients. The accuracy of the estimation was assessed by comparing the results from our ODACoR algorithms with the estimators derived from the meta-analysis and the pooled data. RESULTS: The meta-analysis estimator showed a high relative bias (∼40%) when the clinical condition is relatively rare (∼0.5%), whereas ODACoR algorithms exhibited a substantially lower relative bias (∼0.2%). The estimated effects from our ODACoR algorithms were identical on par with the estimates from the pooled data, suggesting the high reliability of our federated learning algorithms. In contrast, the meta-analysis estimate failed to identify risk factors such as age, gender, chronic conditions history, and obesity, compared to the pooled data. DISCUSSION: Our proposed ODACoR algorithms are communication-efficient, highly accurate, and suitable to characterize the complex interplay between multiple clinical conditions. CONCLUSION: Our study demonstrates that our ODACoR algorithms are communication-efficient and can be widely applicable for analyzing multiple clinical conditions in a time-to-event analysis framework. Dazheng Zhang, Jiayi Tong, Naimin Jing, Chongliang Luo, Dimitri A. Christakis, Diana Güthe, Mady Hornig, Kelly J. Kelleher, Keith E. Morse, Colin M. Rogerson, Jasmin Divers, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 15 |
| 2024 | One-shot distributed algorithms for addressing heterogeneity in competing risks data across clinical sites
Dazheng Zhang, Jiayi Tong, Ronen Stein, Naimin Jing, Mary Regina Boland, Chongliang Luo, Robert N. Baldassano, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Biomed. Informatics | 11 |
| 2023 | Missing data matter: an empirical evaluation of the impacts of missing EHR data in comparative effectiveness researchabstractOBJECTIVES: The impacts of missing data in comparative effectiveness research (CER) using electronic health records (EHRs) may vary depending on the type and pattern of missing data. In this study, we aimed to quantify these impacts and compare the performance of different imputation methods. MATERIALS AND METHODS: We conducted an empirical (simulation) study to quantify the bias and power loss in estimating treatment effects in CER using EHR data. We considered various missing scenarios and used the propensity scores to control for confounding. We compared the performance of the multiple imputation and spline smoothing methods to handle missing data. RESULTS: When missing data depended on the stochastic progression of disease and medical practice patterns, the spline smoothing method produced results that were close to those obtained when there were no missing data. Compared to multiple imputation, the spline smoothing generally performed similarly or better, with smaller estimation bias and less power loss. The multiple imputation can still reduce study bias and power loss in some restrictive scenarios, eg, when missing data did not depend on the stochastic process of disease progression. DISCUSSION AND CONCLUSION: Missing data in EHRs could lead to biased estimates of treatment effects and false negative findings in CER even after missing data were imputed. It is important to leverage the temporal information of disease trajectory to impute missing values when using EHRs as a data resource for CER and to consider the missing rate and the effect size when choosing an imputation method. Yizhao Zhou, Jiasheng Shi, Ronen Stein, Robert N. Baldassano, Christopher B. Forrest, Yong Chen 0016, Jing Huang 0021 |
J. Am. Medical Informatics Assoc. | 6 |
| 2020 | Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithmabstractOBJECTIVES: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites. MATERIALS AND METHODS: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard). RESULTS: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was <3%, and the ratio of standard errors was <1.25 for all scenarios. ODAL2 achieved higher accuracy (with relative bias <0.1% and ratio of standard errors <1.05). In real data analysis, we investigated the associations of 100 medications with fetal loss during pregnancy. We found that ODAL1 provided estimates with relative bias <10% for 85% of medications, and ODAL2 has relative bias <10% for 99% of medications. For communication cost, ODAL1 requires transferring p numbers from each site to the local site and ODAL2 requires transferring (p×p+p) numbers from each site to the local site, where p is the number of parameters in the regression model. CONCLUSIONS: This study demonstrates that ODAL is privacy-preserving and communication-efficient with small bias and high statistical efficiency. Rui Duan 0004, Mary Regina Boland, Howard H. Chang, Hua Xu 0001, Haitao Chu, Christopher H. Schmid, Christopher B. Forrest, John H. Holmes, Martijn J. Schuemie, Jesse A. Berlin, Jason H. Moore, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 9 |
| 2020 | Learning from local to global: An efficient distributed algorithm for modeling time-to-event dataabstractOBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner. Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 11 |
| 2019 | Understanding Early Childhood Obesity via Interpretation of Machine Learning Model PredictionsabstractObesity, as an independent risk factor for increased morbidity and mortality throughout the lifecycle, is a major health issue in the United States. Pediatric obesity is a strong risk factor for adult obesity, as it tends to be stable and tracks into adulthood. Therefore, prevention of childhood obesity is urgently required for reduction in obesity prevalence and obesity related comorbidities. In this paper, the general pediatric obesity development pattern and the onset time period of early childhood obesity was identified via analysis of approximately 11 million pediatric clinical encounters of 860,510 unique individuals. XGBoost model was developed to predict at age 2 years if individuals would develop obesity in early childhood. The model is generalized to both males and females, and achieved an AUC of 81% (± 0.1%). Obesity associated risk factors were further analyzed via interpretation of the XGBoost model predictions. Besides known predictive factors such as weight, height, race, and ethnicity, new factors such as body temperature and respiratory rate were also identified. As body temperature and respiratory rate are related to human metabolism, novel physiologic mechanisms that cause these associations might be discovered in future research. We decomposed model recall to different age ranges when obesity incidence occurred. The model recall for individuals with obesity incidence between 24-36 months was 97.63%, while recall for obesity incidence between 72-84 months was 48.96%, suggesting obesity is less predictable further in the future. Since obesity is largely affected by evolving factors such as life style, diet, and living environment, it is possible that obesity prevention may be achieved via changes in adjustable factors. Xueqin Pang, Christopher B. Forrest, Félice Lê-Scherban, Aaron J. Masino |
ICMLA | 2 |
| 2017 | EHR-based Quality Measurement to Reduce Antibiotic Use in Children
L. Charles Bailey, Hanieh Razzaghi, Elizabeth R. Earley, Jeanhee Moon, Levon Utidjian, Jessica Hawkins, Christopher B. Forrest |
AMIA | 7 |
| 2017 | Developing Computable Phenotypes of Pediatric Chronic Conditions in PEDSnet
Levon Utidjian, Ritu Khare, Hanieh Razzaghi, Amanda F. Dempsey, Michelle Denburg, Christopher B. Forrest, L. Charles Bailey |
AMIA | 6 |
| 2015 | Informatics to support the IOM social and behavioral domains and measuresabstractConsistent collection and use of social and behavioral determinants of health can improve clinical care, prevention and general health, patient satisfaction, research, and public health. A recent Institute of Medicine committee defined a panel of 11 domains and 12 measures to be included in electronic health records. Incorporating the panel into practice creates a number of informatics research opportunities as well as challenges. The informatics issues revolve around standardization, efficient collection and review, decision support, and support for research. The informatics community can aid the effort by simultaneously optimizing the collection of the selected measures while also partnering with social science researchers to develop and validate new sources of information about social and behavioral determinants of health. George Hripcsak, Christopher B. Forrest, Patricia Flatley Brennan, William W. Stead |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Brief communication: PEDSnet: a National Pediatric Learning Health SystemabstractA learning health system (LHS) integrates research done in routine care settings, structured data capture during every encounter, and quality improvement processes to rapidly implement advances in new knowledge, all with active and meaningful patient participation. While disease-specific pediatric LHSs have shown tremendous impact on improved clinical outcomes, a national digital architecture to rapidly implement LHSs across multiple pediatric conditions does not exist. PEDSnet is a clinical data research network that provides the infrastructure to support a national pediatric LHS. A consortium consisting of PEDSnet, which includes eight academic medical centers, two existing disease-specific pediatric networks, and two national data partners form the initial partners in the National Pediatric Learning Health System (NPLHS). PEDSnet is implementing a flexible dual data architecture that incorporates two widely used data models and national terminology standards to support multi-institutional data integration, cohort discovery, and advanced analytics that enable rapid learning. Christopher B. Forrest, Peter A. Margolis, L. Charles Bailey, Keith Marsolo, Mark A. Del Beccaro, Jonathan A. Finkelstein, David E. Milov, Veronica J. Vieland, Bryan A. Wolf, Feliciano B. Yu, Michael G. Kahn |
J. Am. Medical Informatics Assoc. | 1 |
| 2011 | Electronic medical record use in pediatric primary careabstractOBJECTIVES: To characterize patterns of electronic medical record (EMR) use at pediatric primary care acute visits. DESIGN: Direct observational study of 529 acute visits with 27 experienced pediatric clinician users. MEASUREMENTS: For each 20 s interval and at each stage of the visit according to the Davis Observation Code, we recorded whether the physician was communicating with the family only, using the computer while communicating, or using the computer without communication. Regression models assessed the impact of clinician, patient and visit characteristics on overall visit length, time spent interacting with families, and time spent using the computer while interacting. RESULTS: The mean overall visit length was 11:30 (min:sec) with 9:06 spent in the exam room. Clinicians used the EMR during 27% of exam room time and at all stages of the visit (interacting, chatting, and building rapport; history taking; formulation of the diagnosis and treatment plan; and discussing prevention) except the physical exam. Communication with the family accompanied 70% of EMR use. In regression models, computer documentation outside the exam room was associated with visits that were 11% longer (p=0.001), and female clinicians spent more time using the computer while communicating (p=0.003). LIMITATIONS: The 12 study practices shared one EMR. CONCLUSIONS: Among pediatric clinicians with EMR experience, conversation accompanies most EMR use. Our results suggest that efforts to improve EMR usability and clinician EMR training should focus on use in the context of doctor-patient communication. Further study of the impact of documentation inside versus outside the exam room on productivity is warranted. Alexander G. Fiks, Evaline A. Alessandrini, Christopher B. Forrest, Saira Khan, A. Russell Localio, Andreas Gerber |
J. Am. Medical Informatics Assoc. | 3 |