Henry H. Ong

dblp:116/2059 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0001-7960-7363ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021
YearPublicationVenuePosition
2024 Developing and evaluating pediatric phecodes (Peds-Phecodes) for high-throughput phenotyping using electronic health records
abstract
OBJECTIVE: Pediatric patients have different diseases and outcomes than adults; however, existing phecodes do not capture the distinctive pediatric spectrum of disease. We aim to develop specialized pediatric phecodes (Peds-Phecodes) to enable efficient, large-scale phenotypic analyses of pediatric patients. MATERIALS AND METHODS: We adopted a hybrid data- and knowledge-driven approach leveraging electronic health records (EHRs) and genetic data from Vanderbilt University Medical Center to modify the most recent version of phecodes to better capture pediatric phenotypes. First, we compared the prevalence of patient diagnoses in pediatric and adult populations to identify disease phenotypes differentially affecting children and adults. We then used clinical domain knowledge to remove phecodes representing phenotypes unlikely to affect pediatric patients and create new phecodes for phenotypes relevant to the pediatric population. We further compared phenome-wide association study (PheWAS) outcomes replicating known pediatric genotype-phenotype associations between Peds-Phecodes and phecodes. RESULTS: The Peds-Phecodes aggregate 15 533 ICD-9-CM codes and 82 949 ICD-10-CM codes into 2051 distinct phecodes. Peds-Phecodes replicated more known pediatric genotype-phenotype associations than phecodes (248 vs 192 out of 687 SNPs, P < .001). DISCUSSION: We introduce Peds-Phecodes, a high-throughput EHR phenotyping tool tailored for use in pediatric populations. We successfully validated the Peds-Phecodes using genetic replication studies. Our findings also reveal the potential use of Peds-Phecodes in detecting novel genotype-phenotype associations for pediatric conditions. We expect that Peds-Phecodes will facilitate large-scale phenomic and genomic analyses in pediatric populations. CONCLUSION: Peds-Phecodes capture higher-quality pediatric phenotypes and deliver superior PheWAS outcomes compared to phecodes.
Monika E. Grabowska, Sara L. Van Driest, Jamie R. Robinson, Anna E. Patrick, Chris Guardo, Srushti Gangireddy, Henry H. Ong, QiPing Feng, Robert J. Carroll, Prince J. Kannankeril, Wei-Qi Wei
J. Am. Medical Informatics Assoc.7
2024 Large language models facilitate the generation of electronic health record phenotyping algorithms
abstract
OBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms.
Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei
J. Am. Medical Informatics Assoc.2
2023 Evaluating resources composing the PheMAP knowledge base to enhance high-throughput phenotyping
abstract
OBJECTIVE: A previous study, PheMAP, combined independent, online resources to enable high-throughput phenotyping (HTP) using electronic health records (EHRs). However, online resources offer distinct quality descriptions of diseases which may affect phenotyping performance. We aimed to evaluate the phenotyping performance of single resource-based PheMAPs and investigate an optimized strategy for HTP. MATERIALS AND METHODS: We compared how each resource produced top-ranked concept unique identifiers (CUIs) by term frequency-inverse document frequency with Jaccard matrices comparing single resources and the original PheMAP. We correlated top-ranked concepts from each resource to features used in established Phenotype KnowledgeBase (PheKB) algorithms for hypothyroidism, type II diabetes mellitus (T2DM), and dementias. Using resources separately, we calculated multiple phenotype risk scores for individuals from Vanderbilt University Medical Center's BioVU DNA Biobank and compared phenotyping performance against rule-based eMERGE algorithms. Lastly, we implemented an ensemble strategy which classified patient case/control status based upon PheMAP resource agreement. RESULTS: Jaccard similarity matrices indicate that the similarity of CUIs comprising single resource-based PheMAPs varies. Single resource-based PheMAPs generated from MedlinePlus and MedicineNet outperformed others but only encompass 81.6% of overall disease phenotypes. We propose the PheMAP-Ensemble which provides higher average accuracy and precision than the combined average accuracy and precision of single resource-based PheMAPs. While offering complete phenotype coverage, PheMAP-Ensemble significantly increases phenotyping recall compared to the original iteration. CONCLUSIONS: Resources comprising the PheMAP produce different phenotyping performance when implemented individually. The ensemble method significantly improves the quality of PheMAP by fully utilizing dissimilar resources to capture accurate phenotyping data from EHRs.
Nicholas C. Wan, Ali A Yaqoob, Henry H. Ong, Juan Zhao 0003, Wei-Qi Wei
J. Am. Medical Informatics Assoc.3
2023 Evaluating and mitigating bias in machine learning models for cardiovascular disease prediction
Fuchen Li, Patrick Wu, Henry H. Ong, Josh F. Peterson, Wei-Qi Wei, Juan Zhao 0003
J. Biomed. Informatics3
2022 Evaluating Phenotype Classification Using Synthesized Online Content
Wei-Qi Wei, Juan Zhao 0003, Henry H. Ong
AMIA5
2019 Development of a Genomic Data Flow Framework: Results of a Survey Administered to NIH-NHGRI IGNITE and eMERGE Consortia Participants
Paul Richard Dexter, Henry H. Ong, Amanda Elsey, Gillian Bell, Nephi Walton, Wendy K. Chung, Luke V. Rasmussen, J. Kevin Hicks, Aniwaa Owusu-obeng, Stuart A. Scott, Stephen B. Ellis, Josh F. Peterson
AMIA2
2019 Extracting Drug Exposure Epochs and Drug Response Outcomes from Electronic Health Records
Andrea H. Ramirez, Yaping Shi, Elliot M. Fielstein, Jonathan S. Schildcrout, Henry H. Ong, Joshua C. Denny, Josh F. Peterson
AMIA5
2018 Development and Implementation of Genomic Data Pipelines within US institutions: A pilot multi-site survey among NHGRI's IGNITE Genomic Medicine Sites
Josh F. Peterson, Henry H. Ong, Julie A. Lynch, J. Kevin Hicks, Paul Richard Dexter
AMIA2
2018 EHR Extraction of Longitudinal Exposure to Proton Pump Inhibitors
Andrea H. Ramirez, Elliot M. Fielstein, QiPing Feng, Henry H. Ong, Jonathan S. Schildcrout, Yaping Shi, Joshua C. Denny, Josh F. Peterson
AMIA4