EDBT 2026 Demo / reviewers in the wild / expert
Vern Eric Kerchberger
dblp:250/7110
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-0342-1965ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Phecodes: leveraging PheMAP to identify patients lacking diagnosis codes in electronic health recordsabstractOBJECTIVE: Diagnosis codes documented in electronic health records (EHR) are often relied upon to clinically phenotype patients for biomedical research. However, these diagnoses can be incomplete and inaccurate, leading to false negatives when searching for patients with phenotypes of interest. This study aims to determine whether PheMAP, a comprehensive knowledgebase integrating multiple clinical terminologies beyond diagnosis to capture phenotypes, can effectively identify patients lacking relevant EHR diagnosis codes. MATERIALS AND METHODS: We investigated a collection of 3.5 million patient records from Vanderbilt University Medical Center's EHR and focused on 4 well-studied phenotypes: (1) type 2 diabetes mellitus (T2DM), (2) dementia, (3) prostate cancer, and (4) sensorineural hearing loss. We applied PheMAP to match structured concepts in patient records and calculated a phenotype risk score (PheScore) to indicate patient-phenotype similarity. Patients meeting predefined PheScore criteria but lacking diagnosis codes were identified. Clinically knowledgeable experts adjudicated randomly selected patients per phenotype as Positive, Possibly Positive, or Negative. RESULTS: Our approach indicated that 5.3% of patients lacked a diagnosis for T2DM, 4.5% for dementia, 2.2% for prostate cancer, and 0.2% for sensorineural hearing loss. The expert review indicated 100% precision (for Possibly Positive or Positive cases) for dementia and sensorineural hearing loss, and 90.0% and 85.0% precision for T2DM and prostate cancer, respectively. Excluding Possibly Positive cases, the precision for T2DM and prostate cancer was 88.9% and 81.3%, respectively. CONCLUSIONS: Leveraging clinical terminologies incorporated by PheMAP can effectively identify patients with phenotypes who lack EHR diagnosis codes, thereby enhancing phenotyping quality and related research reliability. Chao Yan 0004, Monika E. Grabowska, Rut Thakkar, Alyson L. Dickson, Peter J. Embí, QiPing Feng, Joshua C. Denny, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 8 |
| 2024 | Large language models facilitate the generation of electronic health record phenotyping algorithmsabstractOBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms. Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | Scanning the medical phenome to identify new diagnoses after recovery from COVID-19 in a US cohortabstractOBJECTIVE: COVID-19 survivors are at risk for long-term health effects, but assessing the sequelae of COVID-19 at large scales is challenging. High-throughput methods to efficiently identify new medical problems arising after acute medical events using the electronic health record (EHR) could improve surveillance for long-term consequences of acute medical problems like COVID-19. MATERIALS AND METHODS: We augmented an existing high-throughput phenotyping method (PheWAS) to identify new diagnoses occurring after an acute temporal event in the EHR. We then used the temporal-informed phenotypes to assess development of new medical problems among COVID-19 survivors enrolled in an EHR cohort of adults tested for COVID-19 at Vanderbilt University Medical Center. RESULTS: The study cohort included 186 105 adults tested for COVID-19 from March 5, 2020 to November 1, 2021; of which 30 088 (16.2%) tested positive. Median follow-up after testing was 412 days (IQR 274-528). Our temporal-informed phenotyping was able to distinguish phenotype chapters based on chronicity of their constituent diagnoses. PheWAS with temporal-informed phenotypes identified increased risk for 43 diagnoses among COVID-19 survivors during outpatient follow-up, including multiple new respiratory, cardiovascular, neurological, and pregnancy-related conditions. Findings were robust to sensitivity analyses, and several phenotypic associations were supported by changes in outpatient vital signs or laboratory tests from the pretesting to postrecovery period. CONCLUSION: Temporal-informed PheWAS identified new diagnoses affecting multiple organ systems among COVID-19 survivors. These findings can inform future efforts to enable longitudinal health surveillance for survivors of COVID-19 and other acute medical conditions using the EHR. Vern Eric Kerchberger, Josh F. Peterson, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | De-black-boxing health AI: demonstrating reproducible machine learning computable phenotypes using the N3C-RECOVER Long COVID model in the All of Us data repositoryabstractMachine learning (ML)-driven computable phenotypes are among the most challenging to share and reproduce. Despite this difficulty, the urgent public health considerations around Long COVID make it especially important to ensure the rigor and reproducibility of Long COVID phenotyping algorithms such that they can be made available to a broad audience of researchers. As part of the NIH Researching COVID to Enhance Recovery (RECOVER) Initiative, researchers with the National COVID Cohort Collaborative (N3C) devised and trained an ML-based phenotype to identify patients highly probable to have Long COVID. Supported by RECOVER, N3C and NIH's All of Us study partnered to reproduce the output of N3C's trained model in the All of Us data enclave, demonstrating model extensibility in multiple environments. This case study in ML-based phenotype reuse illustrates how open-source software best practices and cross-site collaboration can de-black-box phenotyping algorithms, prevent unnecessary rework, and promote open science in informatics. Emily R. Pfaff, Andrew T. Girvin, Miles Crosskey, Srushti Gangireddy, Hiral Master, Wei-Qi Wei, Vern Eric Kerchberger, Mark G. Weiner, Paul A. Harris, Melissa A. Basford, Chris Lunt, Christopher G. Chute, Richard A. Moffitt, Melissa A. Haendel |
J. Am. Medical Informatics Assoc. | 7 |
| 2021 | Mapping the Read2/CTV3 controlled clinical terminologies to Phecodes in UK Biobank primary care electronic health records: implementation and evaluation
Spiros C. Denaxas, QiPing Feng, Ghazaleh Fatemifar, Lisa Bastarache, Vern Eric Kerchberger, Aroon D. Hingorani, R. Tom Lumbers, Josh F. Peterson, Wei-Qi Wei, Harry Hemingway |
AMIA | 6 |
| 2021 | ConceptWAS: A high-throughput method for early identification of COVID-19 presenting symptoms and characteristics from clinical notes
Juan Zhao 0003, Monika E. Grabowska, Vern Eric Kerchberger, Joshua C. Smith, H. Nur Eken, QiPing Feng, Josh F. Peterson, S. Trent Rosenbloom, Kevin B. Johnson, Wei-Qi Wei |
J. Biomed. Informatics | 3 |
| 2020 | PheMap: a multi-resource knowledge base for high-throughput phenotyping within electronic health recordsabstractOBJECTIVE: Developing algorithms to extract phenotypes from electronic health records (EHRs) can be challenging and time-consuming. We developed PheMap, a high-throughput phenotyping approach that leverages multiple independent, online resources to streamline the phenotyping process within EHRs. MATERIALS AND METHODS: PheMap is a knowledge base of medical concepts with quantified relationships to phenotypes that have been extracted by natural language processing from publicly available resources. PheMap searches EHRs for each phenotype's quantified concepts and uses them to calculate an individual's probability of having this phenotype. We compared PheMap to clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network for type 2 diabetes mellitus (T2DM), dementia, and hypothyroidism using 84 821 individuals from Vanderbilt Univeresity Medical Center's BioVU DNA Biobank. We implemented PheMap-based phenotypes for genome-wide association studies (GWAS) for T2DM, dementia, and hypothyroidism, and phenome-wide association studies (PheWAS) for variants in FTO, HLA-DRB1, and TCF7L2. RESULTS: In this initial iteration, the PheMap knowledge base contains quantified concepts for 841 disease phenotypes. For T2DM, dementia, and hypothyroidism, the accuracy of the PheMap phenotypes were >97% using a 50% threshold and eMERGE case-control status as a reference standard. In the GWAS analyses, PheMap-derived phenotype probabilities replicated 43 of 51 previously reported disease-associated variants for the 3 phenotypes. For 9 of the 11 top associations, PheMap provided an equivalent or more significant P value than eMERGE-based phenotypes. The PheMap-based PheWAS showed comparable or better performance to a traditional phecode-based PheWAS. PheMap is publicly available online. CONCLUSIONS: PheMap significantly streamlines the process of extracting research-quality phenotype information from EHRs, with comparable or better performance to current phenotyping approaches. Neil S. Zheng, QiPing Feng, Vern Eric Kerchberger, Juan Zhao 0003, Todd L. Edwards, Nancy J. Cox, C. Michael Stein, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Detecting time-evolving phenotypic topics via tensor factorization on electronic health records: Cardiovascular disease case study
Juan Zhao 0003, David J. Schlueter, Patrick Wu, Vern Eric Kerchberger, S. Trent Rosenbloom, Quinn Stanton Wells, QiPing Feng, Joshua C. Denny, Wei-Qi Wei |
J. Biomed. Informatics | 5 |