QiPing Feng

dblp:239/0792 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-6213-793XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 since 2021
YearPublicationVenuePosition
2025 Beyond Phecodes: leveraging PheMAP to identify patients lacking diagnosis codes in electronic health records
abstract
OBJECTIVE: Diagnosis codes documented in electronic health records (EHR) are often relied upon to clinically phenotype patients for biomedical research. However, these diagnoses can be incomplete and inaccurate, leading to false negatives when searching for patients with phenotypes of interest. This study aims to determine whether PheMAP, a comprehensive knowledgebase integrating multiple clinical terminologies beyond diagnosis to capture phenotypes, can effectively identify patients lacking relevant EHR diagnosis codes. MATERIALS AND METHODS: We investigated a collection of 3.5 million patient records from Vanderbilt University Medical Center's EHR and focused on 4 well-studied phenotypes: (1) type 2 diabetes mellitus (T2DM), (2) dementia, (3) prostate cancer, and (4) sensorineural hearing loss. We applied PheMAP to match structured concepts in patient records and calculated a phenotype risk score (PheScore) to indicate patient-phenotype similarity. Patients meeting predefined PheScore criteria but lacking diagnosis codes were identified. Clinically knowledgeable experts adjudicated randomly selected patients per phenotype as Positive, Possibly Positive, or Negative. RESULTS: Our approach indicated that 5.3% of patients lacked a diagnosis for T2DM, 4.5% for dementia, 2.2% for prostate cancer, and 0.2% for sensorineural hearing loss. The expert review indicated 100% precision (for Possibly Positive or Positive cases) for dementia and sensorineural hearing loss, and 90.0% and 85.0% precision for T2DM and prostate cancer, respectively. Excluding Possibly Positive cases, the precision for T2DM and prostate cancer was 88.9% and 81.3%, respectively. CONCLUSIONS: Leveraging clinical terminologies incorporated by PheMAP can effectively identify patients with phenotypes who lack EHR diagnosis codes, thereby enhancing phenotyping quality and related research reliability.
Chao Yan 0004, Monika E. Grabowska, Rut Thakkar, Alyson L. Dickson, Peter J. Embí, QiPing Feng, Joshua C. Denny, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei
J. Am. Medical Informatics Assoc.6
2024 Developing and evaluating pediatric phecodes (Peds-Phecodes) for high-throughput phenotyping using electronic health records
abstract
OBJECTIVE: Pediatric patients have different diseases and outcomes than adults; however, existing phecodes do not capture the distinctive pediatric spectrum of disease. We aim to develop specialized pediatric phecodes (Peds-Phecodes) to enable efficient, large-scale phenotypic analyses of pediatric patients. MATERIALS AND METHODS: We adopted a hybrid data- and knowledge-driven approach leveraging electronic health records (EHRs) and genetic data from Vanderbilt University Medical Center to modify the most recent version of phecodes to better capture pediatric phenotypes. First, we compared the prevalence of patient diagnoses in pediatric and adult populations to identify disease phenotypes differentially affecting children and adults. We then used clinical domain knowledge to remove phecodes representing phenotypes unlikely to affect pediatric patients and create new phecodes for phenotypes relevant to the pediatric population. We further compared phenome-wide association study (PheWAS) outcomes replicating known pediatric genotype-phenotype associations between Peds-Phecodes and phecodes. RESULTS: The Peds-Phecodes aggregate 15 533 ICD-9-CM codes and 82 949 ICD-10-CM codes into 2051 distinct phecodes. Peds-Phecodes replicated more known pediatric genotype-phenotype associations than phecodes (248 vs 192 out of 687 SNPs, P < .001). DISCUSSION: We introduce Peds-Phecodes, a high-throughput EHR phenotyping tool tailored for use in pediatric populations. We successfully validated the Peds-Phecodes using genetic replication studies. Our findings also reveal the potential use of Peds-Phecodes in detecting novel genotype-phenotype associations for pediatric conditions. We expect that Peds-Phecodes will facilitate large-scale phenomic and genomic analyses in pediatric populations. CONCLUSION: Peds-Phecodes capture higher-quality pediatric phenotypes and deliver superior PheWAS outcomes compared to phecodes.
Monika E. Grabowska, Sara L. Van Driest, Jamie R. Robinson, Anna E. Patrick, Chris Guardo, Srushti Gangireddy, Henry H. Ong, QiPing Feng, Robert J. Carroll, Prince J. Kannankeril, Wei-Qi Wei
J. Am. Medical Informatics Assoc.8
2024 Large language models facilitate the generation of electronic health record phenotyping algorithms
abstract
OBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms.
Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei
J. Am. Medical Informatics Assoc.8
2021 Mapping the Read2/CTV3 controlled clinical terminologies to Phecodes in UK Biobank primary care electronic health records: implementation and evaluation
Spiros C. Denaxas, QiPing Feng, Ghazaleh Fatemifar, Lisa Bastarache, Vern Eric Kerchberger, Aroon D. Hingorani, R. Tom Lumbers, Josh F. Peterson, Wei-Qi Wei, Harry Hemingway
AMIA3
2021 DDIWAS: High-throughput electronic health record-based screening of drug-drug interactions
abstract
OBJECTIVE: We developed and evaluated Drug-Drug Interaction Wide Association Study (DDIWAS). This novel method detects potential drug-drug interactions (DDIs) by leveraging data from the electronic health record (EHR) allergy list. MATERIALS AND METHODS: To identify potential DDIs, DDIWAS scans for drug pairs that are frequently documented together on the allergy list. Using deidentified medical records, we tested 616 drugs for potential DDIs with simvastatin (a common lipid-lowering drug) and amlodipine (a common blood-pressure lowering drug). We evaluated the performance to rediscover known DDIs using existing knowledge bases and domain expert review. To validate potential novel DDIs, we manually reviewed patient charts and searched the literature. RESULTS: DDIWAS replicated 34 known DDIs. The positive predictive value to detect known DDIs was 0.85 and 0.86 for simvastatin and amlodipine, respectively. DDIWAS also discovered potential novel interactions between simvastatin-hydrochlorothiazide, amlodipine-omeprazole, and amlodipine-valacyclovir. A software package to conduct DDIWAS is publicly available. CONCLUSIONS: In this proof-of-concept study, we demonstrate the value of incorporating information mined from existing allergy lists to detect DDIs in a real-world clinical setting. Since allergy lists are routinely collected in EHRs, DDIWAS has the potential to detect and validate DDI signals across institutions.
Patrick Wu, Scott D. Nelson, Juan Zhao 0003, Cosby A. Stone Jr., QiPing Feng, Qingxia Chen, Eric A. Larson, Bingshan Li, Nancy J. Cox, C. Michael Stein, Elizabeth Phillips, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei
J. Am. Medical Informatics Assoc.5
2021 ConceptWAS: A high-throughput method for early identification of COVID-19 presenting symptoms and characteristics from clinical notes
Juan Zhao 0003, Monika E. Grabowska, Vern Eric Kerchberger, Joshua C. Smith, H. Nur Eken, QiPing Feng, Josh F. Peterson, S. Trent Rosenbloom, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics6
2021 A retrospective approach to evaluating potential adverse outcomes associated with delay of procedures for cardiovascular and cancer-related diagnoses in the context of COVID-19
Neil S. Zheng, Jeremy L. Warner, Travis Osterman, Quinn Stanton Wells, Xiao-Ou Shu, Steve Deppen, Seth J. Karp, Shon Dwyer, QiPing Feng, Nancy J. Cox, Josh F. Peterson, C. Michael Stein, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei
J. Biomed. Informatics9
2020 Detecting National Institutes of Health's funding interests and trends using Machine Learning
Juan Zhao 0003, QiPing Feng, Wei-Qi Wei
AMIA3
2020 PheMap: a multi-resource knowledge base for high-throughput phenotyping within electronic health records
abstract
OBJECTIVE: Developing algorithms to extract phenotypes from electronic health records (EHRs) can be challenging and time-consuming. We developed PheMap, a high-throughput phenotyping approach that leverages multiple independent, online resources to streamline the phenotyping process within EHRs. MATERIALS AND METHODS: PheMap is a knowledge base of medical concepts with quantified relationships to phenotypes that have been extracted by natural language processing from publicly available resources. PheMap searches EHRs for each phenotype's quantified concepts and uses them to calculate an individual's probability of having this phenotype. We compared PheMap to clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network for type 2 diabetes mellitus (T2DM), dementia, and hypothyroidism using 84 821 individuals from Vanderbilt Univeresity Medical Center's BioVU DNA Biobank. We implemented PheMap-based phenotypes for genome-wide association studies (GWAS) for T2DM, dementia, and hypothyroidism, and phenome-wide association studies (PheWAS) for variants in FTO, HLA-DRB1, and TCF7L2. RESULTS: In this initial iteration, the PheMap knowledge base contains quantified concepts for 841 disease phenotypes. For T2DM, dementia, and hypothyroidism, the accuracy of the PheMap phenotypes were >97% using a 50% threshold and eMERGE case-control status as a reference standard. In the GWAS analyses, PheMap-derived phenotype probabilities replicated 43 of 51 previously reported disease-associated variants for the 3 phenotypes. For 9 of the 11 top associations, PheMap provided an equivalent or more significant P value than eMERGE-based phenotypes. The PheMap-based PheWAS showed comparable or better performance to a traditional phecode-based PheWAS. PheMap is publicly available online. CONCLUSIONS: PheMap significantly streamlines the process of extracting research-quality phenotype information from EHRs, with comparable or better performance to current phenotyping approaches.
Neil S. Zheng, QiPing Feng, Vern Eric Kerchberger, Juan Zhao 0003, Todd L. Edwards, Nancy J. Cox, C. Michael Stein, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei
J. Am. Medical Informatics Assoc.2
2019 Combining Publicly-Available and Electronic Health Record Data to Reposition Drugs
Patrick Wu, QiPing Feng, Joshua C. Denny, Wei-Qi Wei
AMIA2
2019 Deep Learning Using Electronic Health Records and Genetic Data to Predict Cardiovascular Diseases
Juan Zhao 0003, QiPing Feng, Patrick Wu, Joshua C. Denny, Wei-Qi Wei
AMIA2
2019 Detecting time-evolving phenotypic topics via tensor factorization on electronic health records: Cardiovascular disease case study
Juan Zhao 0003, David J. Schlueter, Patrick Wu, Vern Eric Kerchberger, S. Trent Rosenbloom, Quinn Stanton Wells, QiPing Feng, Joshua C. Denny, Wei-Qi Wei
J. Biomed. Informatics8
2018 EHR Extraction of Longitudinal Exposure to Proton Pump Inhibitors
Andrea H. Ramirez, Elliot M. Fielstein, QiPing Feng, Henry H. Ong, Jonathan S. Schildcrout, Yaping Shi, Joshua C. Denny, Josh F. Peterson
AMIA3
2018 Using Topic Modeling to Identify Relationship between LPA Variant and Disease Phenotypes
Juan Zhao 0003, QiPing Feng, Patrick Wu, Joshua C. Denny, Wei-Qi Wei
AMIA2