VLDB 2026 Research / reviewers in the wild / expert
Lisa Bastarache
dblp:68/7612
· DBLP profile ↗
39ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 3 first-author · 16 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Theory and practice in biomedical informatics: a framework for discoveryabstractOBJECTIVE: Clarify disciplinary foundations and internal structure of biomedical informatics. METHODS: We analyze BMI's emergence at disciplinary intersections and map its internal structure across 4 domains: theory and practice of knowledge discovery, knowledge representation and reasoning, knowledge architecture, and knowledge-driven transformation. We compare BMI with mathematics, computer science, biostatistics, and biomedical engineering, and illustrate emergent characteristics through a precision medicine example. RESULTS: BMI's distinctive contribution-elucidating the structure of biomedical knowledge and developing methods to discover, preserve, and make knowledge actionable-requires strength across all 4 domains. BMI developed these domains pragmatically: building systems, extracting principles, and formalizing theories. The discipline must now complement empirical approaches with rigorous theoretical work: assessing adequacy of existing theories, identifying gaps, and orchestrating collaborative development. CONCLUSIONS: BMI creates emergent capabilities across disciplines. As biomedicine becomes increasingly complex, BMI must strengthen its theoretical foundations while demonstrating transformative potential of knowledge spanning biological scales and time. William W. Stead, Constantin F. Aliferis, Lisa Bastarache, Nancy M. Lorenzi, William Edward Hammond |
J. Am. Medical Informatics Assoc. | 3 |
| 2025 | Unmet social needs and diverticulitis: a phenotyping algorithm and cross-sectional analysisabstractOBJECTIVE: To validate a phenotyping algorithm for gradations of diverticular disease severity and investigate relationships between unmet social needs and disease severity. MATERIALS AND METHODS: An algorithm was designed in the All of Us Research Program to identify diverticulosis, mild diverticulitis, and operative or recurrent diverticulitis requiring multiple inpatient admissions. This was validated in an independent institution and applied to a cohort in the All of Us Research Program. Distributions of individual-level social barriers were compared across quintiles of an area-level index through fold enrichment of the barrier in the fifth (most deprived) quintile relative to the first (least deprived) quintile. Social needs of food insecurity, housing instability, and care access were included in logistic regression to assess association with disease severity. RESULTS: Across disease severity groups, the phenotyping algorithm had positive predictive values ranging from 0.87 to 0.97 and negative predictive values ranging from 0.97 to 0.99. Unmet social needs were variably distributed when comparing the most to the least deprived quintile of the area-level deprivation index (fold enrichment ranging from 0.53 to 15). Relative to a reference of diverticulosis, an unmet social need was associated with greater odds of operative or recurrent inpatient diverticulitis (OR [95% CI] 1.61 [1.19-2.17]). DISCUSSION: Understanding the landscape of social barriers in disease-specific cohorts may facilitate a targeted approach when addressing these needs in clinical settings. CONCLUSION: Using a validated phenotyping algorithm for diverticular disease severity, unmet social needs were found to be associated with greater severity of diverticulitis presentation. Thomas E. Ueland, Samuel A Younan, Parker T. Evans, Jessica Sims, Megan M. Shroder, Alexander T. Hawkins, Richard Peek, Xinnan Niu, Lisa Bastarache, Jamie R. Robinson |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | Comparison of phenomic profiles in the All of Us Research Program against the US general population and the UK BiobankabstractIMPORTANCE: Knowledge gained from cohort studies has dramatically advanced both public and precision health. The All of Us Research Program seeks to enroll 1 million diverse participants who share multiple sources of data, providing unique opportunities for research. It is important to understand the phenomic profiles of its participants to conduct research in this cohort. OBJECTIVES: More than 280 000 participants have shared their electronic health records (EHRs) in the All of Us Research Program. We aim to understand the phenomic profiles of this cohort through comparisons with those in the US general population and a well-established nation-wide cohort, UK Biobank, and to test whether association results of selected commonly studied diseases in the All of Us cohort were comparable to those in UK Biobank. MATERIALS AND METHODS: We included participants with EHRs in All of Us and participants with health records from UK Biobank. The estimates of prevalence of diseases in the US general population were obtained from the Global Burden of Diseases (GBD) study. We conducted phenome-wide association studies (PheWAS) of 9 commonly studied diseases in both cohorts. RESULTS: This study included 287 012 participants from the All of Us EHR cohort and 502 477 participants from the UK Biobank. A total of 314 diseases curated by the GBD were evaluated in All of Us, 80.9% (N = 254) of which were more common in All of Us than in the US general population [prevalence ratio (PR) >1.1, P < 2 × 10-5]. Among 2515 diseases and phenotypes evaluated in both All of Us and UK Biobank, 85.6% (N = 2152) were more common in All of Us (PR >1.1, P < 2 × 10-5). The Pearson correlation coefficients of effect sizes from PheWAS between All of Us and UK Biobank were 0.61, 0.50, 0.60, 0.57, 0.40, 0.53, 0.46, 0.47, and 0.24 for ischemic heart diseases, lung cancer, chronic obstructive pulmonary disease, dementia, colorectal cancer, lower back pain, multiple sclerosis, lupus, and cystic fibrosis, respectively. DISCUSSION: Despite the differences in prevalence of diseases in All of Us compared to the US general population or the UK Biobank, our study supports that All of Us can facilitate rapid investigation of a broad range of diseases. CONCLUSION: Most diseases were more common in All of Us than in the general US population or the UK Biobank. Results of disease-disease association tests from All of Us are comparable to those estimated in another well-studied national cohort. Chenjie Zeng, David J. Schlueter, Tam C. Tran, Anav Babbar, Thomas Cassini, Lisa Bastarache, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Disentangling the phenotypic patterns of hypertension and chronic hypotensionabstractOBJECTIVE: 2017 blood pressure (BP) categories focus on cardiac risk. We hypothesize that studying the balance between mechanisms that increase or decrease BP across the medical phenome will lead to new insights. We devised a classifier that uses BP measures to assign individuals to mutually exclusive categories centered in the upper (Htn), lower (Hotn) and middle (Naf) zones of the BP spectrum; and examined the epidemiologic and phenotypic patterns of these BP-categories. METHODS: We classified a cohort of 832,560 deidentified electronic health records by BP-category; compared the frequency of BP-categories and four subtypes of Htn and Hotn by sex and age-decade; visualized the distributions of systolic, diastolic, mean arterial and pulse pressures stratified by BP-category; and ran Phenome-wide Association Studies (PheWAS) for Htn and Hotn. We paired knowledgebases for hypertension and hypotension and computed aggregate knowledgebase status (KB-status) indicating known associations. We assessed alignment of PheWAS results with KB-status for phecodes in the knowledgebase, and paired PheWAS correlations with KB-status to surface phenotypic patterns. RESULTS: BP-categories represent distinct distributions within the multimodal distributions of systolic and diastolic pressure. They are centered in the upper, lower, and middle zones of mean arterial pressure and provide a different signal than pulse pressure. For phecodes in the knowledgebase, 85% of positive correlations align with KB-status. Phenotypic patterns for Htn and Hotn overlap for several phecodes and are separate for others. Our analysis suggests five candidates for hypothesis testing research, two where the prevalence of the association with Htn or Hotn may be under appreciated, three where mechanisms that increase and decrease blood pressure may be affecting one another's expression. CONCLUSION: PairedPheWAS methods may open a phenome-wide path to disentangling hypertension and chronic hypotension. Our classifier provides a starting point for assigning individuals to BP-categories representing the upper, lower, and middle zones of the BP spectrum. 4.7 % of individuals matching 2017 BP categories for normal, elevated BP or isolated hypertension, have diastolic pressure < 60. Research is needed to fine-tune the classifier, provide external validation, evaluate the clinical significance of diastolic pressure < 60, and test the candidate hypotheses. William W. Stead, Adam Lewis, Nunzia Bettinsoli Giuse, Annette M. Williams, Italo Biaggioni, Lisa Bastarache |
J. Biomed. Informatics | 6 |
| 2023 | Next-generation phenotyping: introducing phecodeX for enhanced discovery research in medical phenomicsabstractMOTIVATION: Phecodes are widely used and easily adapted phenotypes based on International Classification of Diseases codes. The current version of phecodes (v1.2) was designed primarily to study common/complex diseases diagnosed in adults; however, there are numerous limitations in the codes and their structure. RESULTS: Here, we present phecodeX, an expanded version of phecodes with a revised structure and 1,761 new codes. PhecodeX adds granularity to phenotypes in key disease domains that are under-represented in the current phecode structure-including infectious disease, pregnancy, congenital anomalies, and neonatology-and is a more robust representation of the medical phenome for global use in discovery research. AVAILABILITY AND IMPLEMENTATION: phecodeX is available at https://github.com/PheWAS/phecodeX. Megan M. Shuey, William W. Stead, Ida Aka, April L. Barnado, Lisa Bastarache, Elly Brokamp, Meredith Campbell, Robert J. Carroll, Jeffrey A. Goldstein, Adam Lewis, Beth A. Malow, Jonathan D. Mosley, Travis Osterman, Dolly A Padovani-Claudio, Andrea Ramirez, Dan M. Roden, Bryce A. Schuler, Edward Siew, Jennifer Sucre, Isaac Thomsen, Rory J. Tinker, Sara Van Driest, Colin Walsh, Jeremy L. Warner, Quinn Stanton Wells, Lee E. Wheless |
Bioinform. | 5 |
| 2023 | Systematic replication of smoking disease associations using survey responses and EHR data in the All of Us Research ProgramabstractOBJECTIVE: The All of Us Research Program (All of Us) aims to recruit over a million participants to further precision medicine. Essential to the verification of biobanks is a replication of known associations to establish validity. Here, we evaluated how well All of Us data replicated known cigarette smoking associations. MATERIALS AND METHODS: We defined smoking exposure as follows: (1) an EHR Smoking exposure that used International Classification of Disease codes; (2) participant provided information (PPI) Ever Smoking; and, (3) PPI Current Smoking, both from the lifestyle survey. We performed a phenome-wide association study (PheWAS) for each smoking exposure measurement type. For each, we compared the effect sizes derived from the PheWAS to published meta-analyses that studied cigarette smoking from PubMed. We defined two levels of replication of meta-analyses: (1) nominally replicated: which required agreement of direction of effect size, and (2) fully replicated: which required overlap of confidence intervals. RESULTS: PheWASes with EHR Smoking, PPI Ever Smoking, and PPI Current Smoking revealed 736, 492, and 639 phenome-wide significant associations, respectively. We identified 165 meta-analyses representing 99 distinct phenotypes that could be matched to EHR phenotypes. At P < .05, 74 were nominally replicated and 55 were fully replicated. At P < 2.68 × 10-5 (Bonferroni threshold), 58 were nominally replicated and 40 were fully replicated. DISCUSSION: Most phenotypes found in published meta-analyses associated with smoking were nominally replicated in All of Us. Both survey and EHR definitions for smoking produced similar results. CONCLUSION: This study demonstrated the feasibility of studying common exposures using All of Us data. David J. Schlueter, Lina M. Sulieman, Huan Mo, Jacob M. Keaton, Tracey Ferrara, Ariel Williams, Onajia J. Stubblefield, Chenjie Zeng, Tam C. Tran, Lisa Bastarache, Anav Babbar, Andrea H. Ramirez, Slavina Goleva, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 11 |
| 2023 | Knowledgebase strategies to aid interpretation of clinical correlation researchabstractOBJECTIVE: Knowledgebases are needed to clarify correlations observed in real-world electronic health record (EHR) data. We posit design principles, present a unifying framework, and report a test of concept. MATERIALS AND METHODS: We structured a knowledge framework along 3 axes: condition of interest, knowledge source, and taxonomy. In our test of concept, we used hypertension as our condition of interest, literature and VanderbiltDDx knowledgebase as sources, and phecodes as our taxonomy. In a cohort of 832 566 deidentified EHRs, we modeled blood pressure and heart rate by sex and age, classified individuals by hypertensive status, and ran a Phenome-wide Association Study (PheWAS) for hypertension. We compared the correlations from PheWAS to the associations in our knowledgebase. RESULTS: We produced PhecodeKbHtn: a knowledgebase comprising 167 hypertension-associated diseases, 15 of which were also negatively associated with blood pressure (pos+neg). Our hypertension PheWAS included 1914 phecodes, 129 of which were in the PhecodeKbHtn. Among the PheWAS association results, phecodes that were in PhecodeKbHtn had larger effect sizes compared with those phecodes not in the knowledgebase. DISCUSSION: Each source contributed unique and additive associations. Models of blood pressure and heart rate by age and sex were consistent with prior cohort studies. All but 4 PheWAS positive and negative correlations for phecodes in PhecodeKbHtn may be explained by knowledgebase associations, hypertensive cardiac complications, or causes of hypertension independently associated with hypotension. CONCLUSION: It is feasible to assemble a knowledgebase that is compatible with EHR data to aid interpretation of clinical correlation research. William W. Stead, Adam Lewis, Nunzia Bettinsoli Giuse, Taneya Y. Koonce, Lisa Bastarache |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | The PheRS R package: Phenotype risk score generation and analysis tool using electronic health record data
Layla Aref, Lisa Bastarache, Jake Hughey |
AMIA | 2 |
| 2022 | The phers R package: using phenotype risk scores based on electronic health records to study Mendelian disease and rare genetic variantsabstractSUMMARY: Electronic health record (EHR) data linked to DNA biobanks are a valuable resource for understanding the phenotypic effects of human genetic variation. We previously developed the phenotype risk score (PheRS) as an approach to quantify the extent to which a patient's clinical features resemble a given Mendelian disease. Using PheRS, we have uncovered novel associations between Mendelian disease-like phenotypes and rare genetic variants, and identified patients who may have undiagnosed Mendelian disease. Although the PheRS approach is conceptually simple, it involves multiple mapping steps and was previously only available as custom scripts, limiting the approach's usability. Thus, we developed the phers R package, a complete and user-friendly set of functions and maps for performing a PheRS-based analysis on linked clinical and genetic data. The package includes up-to-date maps between EHR-based phenotypes (i.e. ICD codes and phecodes), human phenotype ontology terms and Mendelian diseases. Starting with occurrences of ICD codes, the package enables the user to calculate PheRSs, validate the scores using case-control analyses, and perform genetic association analyses. By increasing PheRS's transparency and usability, the phers R package will help improve our understanding of the relationships between rare genetic variants and clinically meaningful human phenotypes. AVAILABILITY AND IMPLEMENTATION: The phers R package is free and open-source and available on CRAN and at https://phers.hugheylab.org. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Layla Aref, Lisa Bastarache, Jacob J. Hughey |
Bioinform. | 2 |
| 2022 | Cox regression is robust to inaccurate EHR-extracted event time: an application to EHR-based GWASabstractMOTIVATION: Logistic regression models are used in genomic studies to analyze the genetic data linked to electronic health records (EHRs), and do not take full usage of the time-to-event information available in EHRs. Previous work has shown that Cox regression, which can account for left truncation and right censoring in EHRs, increased the power to detect genotype-phenotype associations compared to logistic regression. We extend this to evaluate the relative performance of Cox regression and various logistic regression models in the presence of positive errors in event time (delayed event time), relating to recorded event time accuracy. RESULTS: One Cox model and three logistic regression models were considered under different scenarios of delayed event time. Extensive simulations and a genomic study application were used to evaluate the impact of delayed event time. While logistic regression does not model the time-to-event directly, various logistic regression models used in the literature were more sensitive to delayed event time than Cox regression. Results highlighted the importance to identify and exclude the patients diagnosed before entry time. Cox regression had similar or modest improvement in statistical power over various logistic regression models at controlled type I error. This was supported by the empirical data, where the Cox models steadily had the highest sensitivity to detect known genotype-phenotype associations under all scenarios of delayed event time. AVAILABILITY AND IMPLEMENTATION: Access to individual-level EHR and genotype data is restricted by the IRB. Simulation code and R script for data process are at: https://github.com/QingxiaCindyChen/CoxRobustEHR.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Rebecca Irlmeier, Jacob J. Hughey, Lisa Bastarache, Joshua C. Denny, Qingxia Chen |
Bioinform. | 3 |
| 2022 | A research agenda to support the development and implementation of genomics-based clinical informatics tools and resourcesabstractOBJECTIVE: The Genomic Medicine Working Group of the National Advisory Council for Human Genome Research virtually hosted its 13th genomic medicine meeting titled "Developing a Clinical Genomic Informatics Research Agenda". The meeting's goal was to articulate a research strategy to develop Genomics-based Clinical Informatics Tools and Resources (GCIT) to improve the detection, treatment, and reporting of genetic disorders in clinical settings. MATERIALS AND METHODS: Experts from government agencies, the private sector, and academia in genomic medicine and clinical informatics were invited to address the meeting's goals. Invitees were also asked to complete a survey to assess important considerations needed to develop a genomic-based clinical informatics research strategy. RESULTS: Outcomes from the meeting included identifying short-term research needs, such as designing and implementing standards-based interfaces between laboratory information systems and electronic health records, as well as long-term projects, such as identifying and addressing barriers related to the establishment and implementation of genomic data exchange systems that, in turn, the research community could help address. DISCUSSION: Discussions centered on identifying gaps and barriers that impede the use of GCIT in genomic medicine. Emergent themes from the meeting included developing an implementation science framework, defining a value proposition for all stakeholders, fostering engagement with patients and partners to develop applications under patient control, promoting the use of relevant clinical workflows in research, and lowering related barriers to regulatory processes. Another key theme was recognizing pervasive biases in data and information systems, algorithms, access, value, and knowledge repositories and identifying ways to resolve them. Ken Wiley, Laura Findley, Madison Goldrich, Teji Rakhra-Burris, Ana Stevens, Pamela Williams, Carol J. Bult, Rex L. Chisholm, Patricia Deverka, Geoffrey S. Ginsburg, Eric D. Green, Gail P. Jarvik, George A. Mensah, Erin Ramos, Mary Relling, Dan M. Roden, Robb Rowley, Gil Alterovitz, Samuel J. Aronson, Lisa Bastarache, James J. Cimino, Erin L. Crowgey, Guilherme Del Fiol, Robert R. Freimuth, Mark A. Hoffman, Janina M. Jeff, Kevin B. Johnson, Kensaku Kawamoto, Subha Madhavan, Eneida A. Mendonça, Lucila Ohno-Machado, Siddharth Pratap, Casey Overby Taylor, Marylyn D. Ritchie, Nephi Walton, Chunhua Weng, Teresa Zayas-Cabán, Teri A. Manolio, Marc S. Williams |
J. Am. Medical Informatics Assoc. | 20 |
| 2021 | Using Genomic Association Replication Rates as an EHR Quality Measure via the Phenotype-Genotype Reference Map (PGRM)
Sarah DeLozier, Josh F. Peterson, Joshua C. Denny, Lisa Bastarache |
AMIA | 4 |
| 2021 | Mapping the Read2/CTV3 controlled clinical terminologies to Phecodes in UK Biobank primary care electronic health records: implementation and evaluation
Spiros C. Denaxas, QiPing Feng, Ghazaleh Fatemifar, Lisa Bastarache, Vern Eric Kerchberger, Aroon D. Hingorani, R. Tom Lumbers, Josh F. Peterson, Wei-Qi Wei, Harry Hemingway |
AMIA | 5 |
| 2021 | Systematic replication of smoking disease associations in the All of Us Research Program
David J. Schlueter, Lina M. Sulieman, Jacob M. Keaton, Tracey Ferrara, Kyle Webb, Ariel Williams, Francis Ratsimbazafy, Lisa Bastarache, Andrea H. Ramirez, Joshua C. Denny |
AMIA | 9 |
| 2021 | Use of electronic health records to support a public health response to the COVID-19 pandemic in the United States: a perspective from 15 academic medical centersabstractOur goal is to summarize the collective experience of 15 organizations in dealing with uncoordinated efforts that result in unnecessary delays in understanding, predicting, preparing for, containing, and mitigating the COVID-19 pandemic in the US. Response efforts involve the collection and analysis of data corresponding to healthcare organizations, public health departments, socioeconomic indicators, as well as additional signals collected directly from individuals and communities. We focused on electronic health record (EHR) data, since EHRs can be leveraged and scaled to improve clinical care, research, and to inform public health decision-making. We outline the current challenges in the data ecosystem and the technology infrastructure that are relevant to COVID-19, as witnessed in our 15 institutions. The infrastructure includes registries and clinical data networks to support population-level analyses. We propose a specific set of strategic next steps to increase interoperability, overall organization, and efficiencies. Subha Madhavan, Lisa Bastarache, Jeffrey S. Brown, Atul J. Butte, David A. Dorr, Peter J. Embí, Charles P. Friedman, Kevin B. Johnson, Jason H. Moore, Isaac S. Kohane, Philip R. O. Payne, Jessica D. Tenenbaum, Mark G. Weiner, Adam B. Wilcox, Lucila Ohno-Machado |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort
Sarah DeLozier, Sarah Bland, Melissa McPheeters, Quinn Stanton Wells, Eric Farber-Eger, Cosmin Adrian Bejan, Daniel Fabbri, S. Trent Rosenbloom, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei, Josh F. Peterson, Lisa Bastarache |
J. Biomed. Informatics | 13 |
| 2020 | Developing a Phenotype Risk Score for Opioid Adverse Events
Leigh Anne Tang, Sarah DeLozier, Lisa Bastarache, Colin Walsh, Joshua C. Denny |
AMIA | 3 |
| 2019 | Improving the phenotype risk score as a scalable approach to identifying patients with Mendelian diseaseabstractOBJECTIVE: The Phenotype Risk Score (PheRS) is a method to detect Mendelian disease patterns using phenotypes from the electronic health record (EHR). We compared the performance of different approaches mapping EHR phenotypes to Mendelian disease features. MATERIALS AND METHODS: PheRS utilizes Mendelian diseases descriptions annotated with Human Phenotype Ontology (HPO) terms. In previous work, we presented a map linking phecodes (based on International Classification of Diseases [ICD]-Ninth Revision) to HPO terms. For this study, we integrated ICD-Tenth Revision codes and lab data. We also created a new map between HPO terms using customized groupings of ICD codes. We compared the performance with cases and controls for 16 Mendelian diseases using 2.5 million de-identified medical records. RESULTS: PheRS effectively distinguished cases from controls for all 15 positive controls and all approaches tested (P < 4 × 1016). Adding lab data led to a statistically significant improvement for 4 of 14 diseases. The custom ICD groupings improved specificity, leading to an average 8% increase for precision at 100 (-2% to 22%). Eight of 10 adults with cystic fibrosis tested had PheRS in the 95th percentile prio to diagnosis. DISCUSSION: Both phecodes and custom ICD groupings were able to detect differences between affected cases and controls at the population level. The ICD map showed better precision for the highest scoring individuals. Adding lab data improved performance at detecting population-level differences. CONCLUSIONS: PheRS is a scalable method to study Mendelian disease at the population level using electronic health record data and can potentially be used to find patients with undiagnosed Mendelian disease. Lisa Bastarache, Jacob J. Hughey, Jeffery A. Goldstein, Julie A. Bastraache, Satya Das, Neil Charles Zaki, Chenjie Zeng, Leigh Anne Tang, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Computable Longitudinal Patient Trajectories
Jeremy L. Warner, Guergana K. Savova, Noémie Elhadad, Lisa Bastarache, David Gotz |
AMIA | 4 |
| 2017 | Sub-Phenotyping of Crohn's Disease Using a Large Electronic Record Cohort
Jamie R. Robinson, Lisa Bastarache, Robert J. Carroll, Elizabeth A. Scoville, David A. Schwartz, Joshua C. Denny |
AMIA | 2 |
| 2017 | Association of BMI and Obesity Genetic Risk Score with Surgical Procedures Through a Procedure-wide Association Study
Jamie R. Robinson, Zongyang Mou, Lisa Bastarache, Wei-Qi Wei, Robert J. Carroll, Joshua C. Denny |
AMIA | 3 |
| 2017 | Evaluating electronic health record data sources and algorithmic approaches to identify hypertensive individualsabstractOBJECTIVE: Phenotyping algorithms applied to electronic health record (EHR) data enable investigators to identify large cohorts for clinical and genomic research. Algorithm development is often iterative, depends on fallible investigator intuition, and is time- and labor-intensive. We developed and evaluated 4 types of phenotyping algorithms and categories of EHR information to identify hypertensive individuals and controls and provide a portable module for implementation at other sites. MATERIALS AND METHODS: We reviewed the EHRs of 631 individuals followed at Vanderbilt for hypertension status. We developed features and phenotyping algorithms of increasing complexity. Input categories included International Classification of Diseases, Ninth Revision (ICD9) codes, medications, vital signs, narrative-text search results, and Unified Medical Language System (UMLS) concepts extracted using natural language processing (NLP). We developed a module and tested portability by replicating 10 of the best-performing algorithms at the Marshfield Clinic. RESULTS: Random forests using billing codes, medications, vitals, and concepts had the best performance with a median area under the receiver operator characteristic curve (AUC) of 0.976. Normalized sums of all 4 categories also performed well (0.959 AUC). The best non-NLP algorithm combined normalized ICD9 codes, medications, and blood pressure readings with a median AUC of 0.948. Blood pressure cutoffs or ICD9 code counts alone had AUCs of 0.854 and 0.908, respectively. Marshfield Clinic results were similar. CONCLUSION: This work shows that billing codes or blood pressure readings alone yield good hypertension classification performance. However, even simple combinations of input categories improve performance. The most complex algorithms classified hypertension with excellent recall and precision. Pedro L. Teixeira, Wei-Qi Wei, Robert M. Cronin, Huan Mo, Jacob P. VanHouten, Robert J. Carroll, Eric LaRose, Lisa Bastarache, S. Trent Rosenbloom, Todd L. Edwards, Dan M. Roden, Thomas A. Lasko, Richard A. Dart, Anne M. Nikolai, Peggy L. Peissig, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 8 |
| 2015 | PheWAS Network Analysis and Visualization
Yaomin Xu, Todd L. Edwards, Lisa Bastarache, Rebecca N. Jerome, Shilin Zho, Eric Torstenson, Wei-Qi Wei, Jana Shirey-Rice, Erica A. Bowton, Shyr Yu, Jill M. Pulley, Joshua C. Denny |
AMIA | 3 |
| 2014 | Phenome-Wide Association Studies Using NLP-Derived Concepts
Pedro L. Teixeira, Robert J. Carroll, Lisa Bastarache, Peter Speltz, Joshua C. Smith, Joshua C. Denny |
AMIA | 3 |
| 2014 | R PheWAS: data analysis and plotting tools for phenome-wide association studies in the R environmentabstractUNLABELLED: Phenome-wide association studies (PheWAS) have been used to replicate known genetic associations and discover new phenotype associations for genetic variants. This PheWAS implementation allows users to translate ICD-9 codes to PheWAS case and control groups, perform analyses using these and/or other phenotypes with covariate adjustments and plot the results. We demonstrate the methods by replicating a PheWAS on rs3135388 (near HLA-DRB, associated with multiple sclerosis) and performing a novel PheWAS using an individual's maximum white blood cell count (WBC) as a continuous measure. Our results for rs3135388 replicate known associations with more significant results than the original study on the same dataset. Our PheWAS of WBC found expected results, including associations with infections, myeloproliferative diseases and associated conditions, such as anemia. These results demonstrate the performance of the improved classification scheme and the flexibility of PheWAS encapsulated in this package. AVAILABILITY AND IMPLEMENTATION: This R package is freely available under the Gnu Public License (GPL-3) from http://phewascatalog.org. It is implemented in native R and is platform independent. Robert J. Carroll, Lisa Bastarache, Joshua C. Denny |
Bioinform. | 2 |
| 2013 | Classifying ICD-9 codes into meaningful disease categories: A comparison between two coding systems
Lisa Bastarache, Wei-Qi Wei, Joshua C. Denny |
AMIA | 1 |
| 2013 | Open Source R Implementation of the PheWAS Methodology
Robert J. Carroll, Lisa Bastarache, Joshua C. Denny |
AMIA | 2 |
| 2013 | A Natural Language Processing Algorithm to define a Venous Thromboembolism Phenotype
Eugenia R. McPeek Hinz, Joshua C. Denny, Lisa Bastarache |
AMIA | 3 |
| 2013 | Using PheWAS and Natural Language Processing to Discover Clinical Associations for Congenital Chest Deformities
Christine M. McEvoy, Robert J. Carroll, Lisa Bastarache, Wei-Qi Wei, Joshua C. Denny |
AMIA | 3 |
| 2013 | Validation and Enhancement of a Computable Medication Indication Resource (MEDI) Using a Large Practice-based Dataset
Wei-Qi Wei, Jonathan D. Mosley, Lisa Bastarache, Joshua C. Denny |
AMIA | 3 |
| 2013 | Development and evaluation of an ensemble resource linking medications to their indicationsabstractOBJECTIVE: To create a computable MEDication Indication resource (MEDI) to support primary and secondary use of electronic medical records (EMRs). MATERIALS AND METHODS: We processed four public medication resources, RxNorm, Side Effect Resource (SIDER) 2, MedlinePlus, and Wikipedia, to create MEDI. We applied natural language processing and ontology relationships to extract indications for prescribable, single-ingredient medication concepts and all ingredient concepts as defined by RxNorm. Indications were coded as Unified Medical Language System (UMLS) concepts and International Classification of Diseases, 9th edition (ICD9) codes. A total of 689 extracted indications were randomly selected for manual review for accuracy using dual-physician review. We identified a subset of medication-indication pairs that optimizes recall while maintaining high precision. RESULTS: MEDI contains 3112 medications and 63 343 medication-indication pairs. Wikipedia was the largest resource, with 2608 medications and 34 911 pairs. For each resource, estimated precision and recall, respectively, were 94% and 20% for RxNorm, 75% and 33% for MedlinePlus, 67% and 31% for SIDER 2, and 56% and 51% for Wikipedia. The MEDI high-precision subset (MEDI-HPS) includes indications found within either RxNorm or at least two of the three other resources. MEDI-HPS contains 13 304 unique indication pairs regarding 2136 medications. The mean±SD number of indications for each medication in MEDI-HPS is 6.22 ± 6.09. The estimated precision of MEDI-HPS is 92%. CONCLUSIONS: MEDI is a publicly available, computable resource that links medications with their indications as represented by concepts and billing codes. MEDI may benefit clinical EMR applications and reuse of EMR data for research. Wei-Qi Wei, Robert M. Cronin, Hua Xu 0001, Thomas A. Lasko, Lisa Bastarache, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 5 |
| 2012 | Diabetes and Susceptibility to Infection: A Study of Lab Culture Results in the EMR
Lisa Bastarache, Wei-Qi Wei, Joshua C. Denny |
AMIA | 1 |
| 2012 | Using PheWAS to Assess Pleiotropy of Genetic Risk Scores for Rheumatoid Arthritis and Coronary Artery Disease in the eMERGE Network
Robert J. Carroll, Katherine P. Liao, Anne E. Eyler, Lisa Bastarache, Dana C. Crawford, Peggy L. Peissig, Jyotishman Pathak, David Carrell, Abel N. Kho, Rongling Li, Daniel R. Masys, Gail P. Jarvik, Christopher G. Chute, Rex L. Chisholm, Eric B. Larson, Catherine A. McCarty, Iftikhar J. Kullo |
AMIA | 4 |
| 2012 | Comparing Diagnoses Recorded in Problem Lists vs. Administrative Codes
Wei-Qi Wei, Lisa Bastarache, Joshua C. Denny |
AMIA | 2 |
| 2010 | PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene-disease associationsabstractMOTIVATION: Emergence of genetic data coupled to longitudinal electronic medical records (EMRs) offers the possibility of phenome-wide association scans (PheWAS) for disease-gene associations. We propose a novel method to scan phenomic data for genetic associations using International Classification of Disease (ICD9) billing codes, which are available in most EMR systems. We have developed a code translation table to automatically define 776 different disease populations and their controls using prevalent ICD9 codes derived from EMR data. As a proof of concept of this algorithm, we genotyped the first 6005 European-Americans accrued into BioVU, Vanderbilt's DNA biobank, at five single nucleotide polymorphisms (SNPs) with previously reported disease associations: atrial fibrillation, Crohn's disease, carotid artery stenosis, coronary artery disease, multiple sclerosis, systemic lupus erythematosus and rheumatoid arthritis. The PheWAS software generated cases and control populations across all ICD9 code groups for each of these five SNPs, and disease-SNP associations were analyzed. The primary outcome of this study was replication of seven previously known SNP-disease associations for these SNPs. RESULTS: Four of seven known SNP-disease associations using the PheWAS algorithm were replicated with P-values between 2.8 x 10(-6) and 0.011. The PheWAS algorithm also identified 19 previously unknown statistical associations between these SNPs and diseases at P < 0.01. This study indicates that PheWAS analysis is a feasible method to investigate SNP-disease associations. Further evaluation is needed to determine the validity of these associations and the appropriate statistical thresholds for clinical significance. AVAILABILITY: The PheWAS software and code translation table are freely available at http://knowledgemap.mc.vanderbilt.edu/research. Joshua C. Denny, Marylyn D. Ritchie, Melissa A. Basford, Jill M. Pulley, Lisa Bastarache, Kristin Brown-Gentry, Deede Wang, Daniel R. Masys, Dan M. Roden, Dana C. Crawford |
Bioinform. | 5 |
| 2010 | Extracting timing and status descriptors for colonoscopy testing from electronic medical recordsabstractColorectal cancer (CRC) screening rates are low despite confirmed benefits. The authors investigated the use of natural language processing (NLP) to identify previous colonoscopy screening in electronic records from a random sample of 200 patients at least 50 years old. The authors developed algorithms to recognize temporal expressions and 'status indicators', such as 'patient refused', or 'test scheduled'. The new methods were added to the existing KnowledgeMap concept identifier system, and the resulting system was used to parse electronic medical records (EMR) to detect completed colonoscopies. Using as the 'gold standard' expert physicians' manual review of EMR notes, the system identified timing references with a recall of 0.91 and precision of 0.95, colonoscopy status indicators with a recall of 0.82 and precision of 0.95, and references to actually completed colonoscopies with recall of 0.93 and precision of 0.95. The system was superior to using colonoscopy billing codes alone. Health services researchers and clinicians may find NLP a useful adjunct to traditional methods to detect CRC screening status. Further investigations must validate extension of NLP approaches for other types of CRC screening applications. Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson |
J. Am. Medical Informatics Assoc. | 6 |
| 2010 | Integrating existing natural language processing tools for medication extraction from discharge summariesabstractOBJECTIVE: To develop an automated system to extract medications and related information from discharge summaries as part of the 2009 i2b2 natural language processing (NLP) challenge. This task required accurate recognition of medication name, dosage, mode, frequency, duration, and reason for drug administration. DESIGN: We developed an integrated system using several existing NLP components developed at Vanderbilt University Medical Center, which included MedEx (to extract medication information), SecTag (a section identification system for clinical notes), a sentence splitter, and a spell checker for drug names. Our goal was to achieve good performance with minimal to no specific training for this document corpus; thus, evaluating the portability of those NLP tools beyond their home institution. The integrated system was developed using 17 notes that were annotated by the organizers and evaluated using 251 notes that were annotated by participating teams. MEASUREMENTS: The i2b2 challenge used standard measures, including precision, recall, and F-measure, to evaluate the performance of participating systems. There were two ways to determine whether an extracted textual finding is correct or not: exact matching or inexact matching. The overall performance for all six types of medication-related findings across 251 annotated notes was considered as the primary metric in the challenge. RESULTS: Our system achieved an overall F-measure of 0.821 for exact matching (0.839 precision; 0.803 recall) and 0.822 for inexact matching (0.866 precision; 0.782 recall). The system ranked second out of 20 participating teams on overall performance at extracting medications and related information. CONCLUSIONS: The results show that the existing MedEx system, together with other NLP components, can extract medication information in clinical text from institutions other than the site of algorithm development with reasonable performance. Son Doan, Lisa Bastarache, Sergio Klimkowski, Joshua C. Denny, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2009 | Development of a Natural Language Processing System to Identify Timing and Status of Colonoscopy Testing in Electronic Medical Records
Joshua C. Denny, Josh F. Peterson, Neesha N. Choma, Hua Xu 0001, Randolph A. Miller, Lisa Bastarache, Neeraja B. Peterson |
AMIA | 6 |
| 2009 | Tracking medical students' clinical experiences using natural language processing
Joshua C. Denny, Lisa Bastarache, Elizabeth Ann Sastre, Anderson Spickard III |
J. Biomed. Informatics | 2 |