L. Charles Bailey

dblp:146/8832 · also Charles Bailey 0001 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0002-8967-0662ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Multi-site analysis of COVID-19 and new-onset diabetes reveals need for improved sensitivity of EHR-based COVID-19 phenotypes - a DiCAYA Network analysis
abstract
OBJECTIVE: We discuss implications of potential ascertainment biases for studies examining diabetes risk following SARS-CoV-2 infection using electronic health records (EHRs). We quantitatively explore sensitivity of results to misclassification of COVID-19 status using data from the U.S.-based Diabetes in Children, Adolescents and Young Adults (DiCAYA) Network on children (≤17 years) and young adults (18-44 years). MATERIALS AND METHODS: In our retrospective case study from the DiCAYA Network, SARS-CoV-2 was identified using labs and diagnoses from June 1, 2020 to December 31, 2021. Patients were followed through December 31, 2022 for new diabetes diagnoses. Sites examined incident diabetes by COVID-19 status using Cox proportional hazards models. Results were pooled in meta-analyses. A bias analysis examined potential impact of COVID-19 misclassification scenarios on results, guided by hypotheses that sensitivity would be <50% and would be higher among those who developed diabetes. RESULTS: Prevalence of documented COVID-19 was low overall and variable across sites (children: 4.4%-7.7%, young adults: 6.2%-22.7%). Individuals with documented COVID-19 were at higher risk of incident diabetes compared to those with no documented infection, but results were heterogeneous across sites. Findings were highly sensitive to COVID-19 misclassification assumptions. Observed results could be biased away from the null under several differential misclassification scenarios. DISCUSSION: Although EHR-based documentation of COVID-19 was associated with incident diabetes, COVID-19 phenotypes likely had low sensitivity, with considerable variation across sites. Misclassification assumptions strongly impacted interpretation of results. CONCLUSION: Given the potential for low phenotype sensitivity and misclassification, caution is warranted when interpreting analyses of COVID-19 and incident diabetes using clinical or administrative databases.
Lorna E. Thorpe, Jasmin Divers, Annemarie Hirsch, Brian S. Schwartz, Jihad S. Obeid, Angela Liese, Tessa L. Crume, Anna Bellatorre, Jiang Bian 0001, Yi Guo 0005, Sarah Bost, Tianchen Lyu, Matthew T. Mefford, Matt Zhou, Eva Lustigova, Levon Utidjian, Mitchell Maltenfort, Patrick Hanley, Meda E. Pavkov, Marc B. Rosenman, Andrea R. Titus, L. Charles Bailey, Christopher B. Forrest, Mitch Maltenfort, Amy Shah, Eneida A. Mendonça, G. Todd Alonso, Sara J. Deakyne Davies, H. Timothy Bunnell, Anne Kazak, Melody Kitzmiller, Manmohan Kamboj, Dimitri A. Christakis, Daksha Ranade, Annemarie G. Hirsch, Joseph J. Dewalle, H. Lester Kirchner, Meredith Lewis, Dione G. Mercer, Cara M. Nordberg, Amy Poissant, Brian E. Dixon, Shaun J. Grannis, Katie Allen, Anna Roberts, Nimish Valvi, Jeff Warvel, Ashley Wiensch, Tamara S. Hannon, Kristi Reynolds, John Chang, Don McCarthy, Rong Wei, Marc Rosenman, George Lales, Anthony Wong, Allison Zelinski, Yuan Luo 0001, Mark Weiner, Pedro Rivera, Thomas Carton, Elizabeth Nauman, Harold P. Lehmann, Meredith Akerman, Rebecca Anthopolos, Stefanie Bendik, Sarah Conderino, Andrew Fair, Jessica Guillaume, Shahidul Islam, Alan Jacobson, David C. Lee, Chinyere Okpara, Anand Rajan, Andrea Titus, Dana Dabelea, Theresa Anderson, Rebecca Conway, Toan Ong, Jack Pattee, Shawna Burgett, Elizabeth Shenkman, William T. Donahoo, William R. Hogan, Piaopiao Li, Mattia Prosperi, Yonghui Wu 0001, Angela D. Liese, Lisa Knight, Caroline Rudisill, Jessica Stucker, Deborah Bowlby, Elaine Apperson, Alex Ewing, Giuseppina Imperatore, Deborah Rolka, Ibrahim Zaganjor
J. Am. Medical Informatics Assoc.23
2026 A multifaceted approach to advancing data quality and fitness standards in multi-institutional networks
abstract
OBJECTIVE: To construct a data quality (DQ) system that incorporates combinations of methods to evaluate data characteristics and analytic fitness across research questions for multiple uses. MATERIALS AND METHODS: Drawing from experience of other data quality programs, network data extraction needs, and recurring study requirements, we developed 5 standards to guide development of a modular, multifaceted data quality system. These included annotation and documentation, ability to measure research readiness, reproducibility across networks, flexibility for the user, and interpretability to research and project teams. Implementation of checks based on these principles focused on reusability and interactive visualization of results. RESULTS: We identified 10 check types producing over 444 check applications and deployed them in 2 multi-institutional networks. Check types span structural conformance to a data model, utility for common research needs, and study-specific customization. All check types are customizable without dependencies between them. A dashboard visualizes results, permitting adjustments based on number of data sources, need for source masking, and the user's focus. All components can be applied as written to any data source using OMOP and are readily modified for other data models. DISCUSSION: We have extended previous work through our novel and multifaceted approach to data quality assessment, addressing needs in both network data improvement and research usage. We developed a capable and deployable system rather than tailoring to specific use cases. CONCLUSION: Our novel DQ assessment system provides essential components for future standardization and collaboration to improve fitness of clinical data for intended use.
Hanieh Razzaghi, Kimberley Dickinson, Kaleigh Wieand, Samuel Boss, Hunter Weidlich, Yungui Huang, Keith E. Morse, Sujan Kumar Mutyala, Jyothi Priya Alekapatti Nandagopal, Karthik Viswanathan, Christopher B. Forrest, L. Charles Bailey
J. Am. Medical Informatics Assoc.12
2023 Assessing the impact of privacy-preserving record linkage on record overlap and patient demographic and clinical characteristics in PCORnet®, the National Patient-Centered Clinical Research Network
abstract
OBJECTIVE: This article describes the implementation of a privacy-preserving record linkage (PPRL) solution across PCORnet®, the National Patient-Centered Clinical Research Network. MATERIAL AND METHODS: Using a PPRL solution from Datavant, we quantified the degree of patient overlap across the network and report a de-duplicated analysis of the demographic and clinical characteristics of the PCORnet population. RESULTS: There were ∼170M patient records across the responding Network Partners, with ∼138M (81%) of those corresponding to a unique patient. 82.1% of patients were found in a single partner and 14.7% were in 2. The percentage overlap between Partners ranged between 0% and 80% with a median of 0%. Linking patients' electronic health records with claims increased disease prevalence in every clinical characteristic, ranging between 63% and 173%. DISCUSSION: The overlap between Partners was variable and depended on timeframe. However, patient data linkage changed the prevalence profile of the PCORnet patient population. CONCLUSIONS: This project was one of the largest linkage efforts of its kind and demonstrates the potential value of record linkage. Linkage between Partners may be most useful in cases where there is geographic proximity between Partners, an expectation that potential linkage Partners will be able to fill gaps in data, or a longer study timeframe.
Keith Marsolo, Daniel Kiernan, Sengwee Toh, Jasmin Phua, Darcy Louzao, Kevin Haynes, Mark G. Weiner, Francisco Angulo, L. Charles Bailey, Jiang Bian 0001, Daniel Fort, Shaun J. Grannis, Ashok K. Krishnamurthy 0001, Vinit Nair, Pedro Rivera, Jonathan C. Silverstein, Maryan Zirkle, Thomas Carton
J. Am. Medical Informatics Assoc.9
2022 The Informatics of RECOVER: Understanding the Post Acute Sequelae of SARS-CoV-2 Infection
Mark G. Weiner, L. Charles Bailey, Richard R. Moffitt, Shawn N. Murphy
AMIA2
2021 U.S. COVID-19 Surveillance in PCORnet®
Sheryl A. Kluberg, Thomas Carton, L. Charles Bailey, Julia A. Fearrington, Keith Marsolo, Kshema M. Nagavedu, Jon Puro, Jason P. Block
AMIA3
2020 Risk Classification in Pediatric Acute Lymphoblastic Leukemia with Structured Data in Clinical Research Networks
Hanieh Razzaghi, Jason Roy, Ashley Batugo, L. Charles Bailey
AMIA4
2020 A Framework for Analysis, Ontological Evaluation, and Visualization in Preparation to Predictive Analytics in Pediatric Brain Tumor Research
abstract
We provide a generalizable framework for the systematic analysis of complicated, longitudinal clinical features in pediatric cancer. We use a threefold pipeline of exploratory data analysis, ontological categorization through a multi-modal data transformation process towards predictive analytics. We derive a data-driven phenotype from a subset of a sample of over 1900 brain tumor cases focused specifically on High-Grade Gliomas. We implement an analyst-friendly process to make machine learning-ready data sets based on domain ontologies ready for enumeration and vectorization. The results are clinical domain expert readable data points from 4.3 million observational events across 16,000 patient days. In this research, we address the gap in phenotypic data features by utilizing extensive harmonized observational clinical data and identify resources and specific processes for their use in rare tumor research.
Alex S. Felmeister, Angela J. Waanders, Jennifer L. Mason, Jeffrey Stevens, L. Charles Bailey, Shiva Ganesan, Ingo Helbig
BIBM5
2018 Probabilistic Linkage of Virtual Pediatric Systems and PEDSnet Patients
Adam C. Dziorny, Robert B. Lindell, L. Charles Bailey
AMIA3
2017 EHR-based Quality Measurement to Reduce Antibiotic Use in Children
L. Charles Bailey, Hanieh Razzaghi, Elizabeth R. Earley, Jeanhee Moon, Levon Utidjian, Jessica Hawkins, Christopher B. Forrest
AMIA1
2017 Standardization of Prescribing Data in PCORnet: RxNorm Concept Unique Identifiers in Multi-Site Research
Casie E. Horgan, Jessica L. Sturtevant, L. Charles Bailey, Lindsey E. Petro, Julia A. Fearrington, Juliane S. Reynolds, Jeffrey S. Brown, Jason P. Block
AMIA3
2017 Developing Computable Phenotypes of Pediatric Chronic Conditions in PEDSnet
Levon Utidjian, Ritu Khare, Hanieh Razzaghi, Amanda F. Dempsey, Michelle Denburg, Christopher B. Forrest, L. Charles Bailey
AMIA7
2017 Preliminary exploratory data analysis of simulated national clinical data research network for future use in annotation of a rare tumor biobanking initiative
abstract
Observational data resources based on the capture of clinical data in the electronic health record (EHR) have produced significant learning opportunities in many areas of medicine. These large data resources can span multiple hospital systems and employ common semantics, ontologies, and data models. They have uncovered critical safety issues for patients, and spurred observational research and clinical decision support. In the age of precision medicine there is also an increased need to obtain genomic and clinical data to discover novel treatments for the deadliest of diseases. With this, there are efforts to create deep-dive disease specific repositories that include tissue in biobanks. The latter require significant human annotation of biospecimens. Securing the data is especially critical in rare pediatric brain tumors. In the specific case of The Children's Brain Tumor Tissue Consortium (CBTTC) an international rare pediatric brain tumor repository, the number of patients that need to be followed prospectively is outpacing the ability of human annotation. In this preliminary study, we perform a prescribed data exploration analysis on simulation data in the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) employed by the pediatric data network PEDSNet with the intention to ascertain feasibility in automatic annotation of patient records in the CBTTC.
Alex S. Felmeister, Angela J. Waanders, Sarah E. S. Leary, Jeffrey Stevens, Jennifer L. Mason, Rachel Teneralli, Xiaohua Hu 0001, L. Charles Bailey
BIBM8
2017 A longitudinal analysis of data quality in a large pediatric data research network
abstract
OBJECTIVE: PEDSnet is a clinical data research network (CDRN) that aggregates electronic health record data from multiple children's hospitals to enable large-scale research. Assessing data quality to ensure suitability for conducting research is a key requirement in PEDSnet. This study presents a range of data quality issues identified over a period of 18 months and interprets them to evaluate the research capacity of PEDSnet. MATERIALS AND METHODS: Results were generated by a semiautomated data quality assessment workflow. Two investigators reviewed programmatic data quality issues and conducted discussions with the data partners' extract-transform-load analysts to determine the cause for each issue. RESULTS: The results include a longitudinal summary of 2182 data quality issues identified across 9 data submission cycles. The metadata from the most recent cycle includes annotations for 850 issues: most frequent types, including missing data (>300) and outliers (>100); most complex domains, including medications (>160) and lab measurements (>140); and primary causes, including source data characteristics (83%) and extract-transform-load errors (9%). DISCUSSION: The longitudinal findings demonstrate the network's evolution from identifying difficulties with aligning the data to a common data model to learning norms in clinical pediatrics and determining research capability. CONCLUSION: While data quality is recognized as a critical aspect in establishing and utilizing a CDRN, the findings from data quality assessments are largely unpublished. This paper presents a real-world account of studying and interpreting data quality findings in a pediatric CDRN, and the lessons learned could be used by other CDRNs.
Ritu Khare, Levon Utidjian, Byron Ruth, Michael G. Kahn, Evanette Burrows, Keith Marsolo, Nandan Patibandla, Hanieh Razzaghi, Ryan Colvin, Daksha Ranade, Melody Kitzmiller, Daniel Eckrich, L. Charles Bailey
J. Am. Medical Informatics Assoc.13
2016 PEDSnet: from building a high-quality CDRN to conducting science
L. Charles Bailey, Michael G. Kahn, Sara J. Deakyne Davies, Ritu Khare, Katherine Deans
AMIA1
2015 Identifying and Understanding Data Quality Issues in a Pediatric Distributed Research Network
Ritu Khare, Levon Utidjian, Gregory Schulte, Keith Marsolo, L. Charles Bailey
AMIA5
2014 Brief communication: PEDSnet: a National Pediatric Learning Health System
abstract
A learning health system (LHS) integrates research done in routine care settings, structured data capture during every encounter, and quality improvement processes to rapidly implement advances in new knowledge, all with active and meaningful patient participation. While disease-specific pediatric LHSs have shown tremendous impact on improved clinical outcomes, a national digital architecture to rapidly implement LHSs across multiple pediatric conditions does not exist. PEDSnet is a clinical data research network that provides the infrastructure to support a national pediatric LHS. A consortium consisting of PEDSnet, which includes eight academic medical centers, two existing disease-specific pediatric networks, and two national data partners form the initial partners in the National Pediatric Learning Health System (NPLHS). PEDSnet is implementing a flexible dual data architecture that incorporates two widely used data models and national terminology standards to support multi-institutional data integration, cohort discovery, and advanced analytics that enable rapid learning.
Christopher B. Forrest, Peter A. Margolis, L. Charles Bailey, Keith Marsolo, Mark A. Del Beccaro, Jonathan A. Finkelstein, David E. Milov, Veronica J. Vieland, Bryan A. Wolf, Feliciano B. Yu, Michael G. Kahn
J. Am. Medical Informatics Assoc.3
2007 Functional dissipation microarrays for classification
Domenico Napoletani, Daniele C. Struppa, Timothy D. Sauer, Victor Morozov, Nikolai Vsevolodov, L. Charles Bailey
Pattern Recognit.6