VLDB 2026 Research / reviewers in the wild / expert
Dan M. Roden
dblp:47/5181
· DBLP profile ↗
28ranked-venue papers
0as first author
11since 2021 · last 2024
0000-0002-6302-0389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 28 · 11 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Empowering the biomedical research community: Innovative SAS deployment on the All of Us Researcher WorkbenchabstractOBJECTIVES: The All of Us Research Program is a precision medicine initiative aimed at establishing a vast, diverse biomedical database accessible through a cloud-based data analysis platform, the Researcher Workbench (RW). Our goal was to empower the research community by co-designing the implementation of SAS in the RW alongside researchers to enable broader use of All of Us data. MATERIALS AND METHODS: Researchers from various fields and with different SAS experience levels participated in co-designing the SAS implementation through user experience interviews. RESULTS: Feedback and lessons learned from user testing informed the final design of the SAS application. DISCUSSION: The co-design approach is critical for reducing technical barriers, broadening All of Us data use, and enhancing the user experience for data analysis on the RW. CONCLUSION: Our co-design approach successfully tailored the implementation of the SAS application to researchers' needs. This approach may inform future software implementations on the RW. Izabelle P. Humes, Cathy Shyr, Moira Dillon, Zhongjie Liu, Jennifer Peterson, Chris De St. Jeor, Jacqueline Malkes, Hiral Master, Brandy Mapes, Romuladus Azuine, Nakia Mack, Bassent Abdelbary, Joyonna Gamble-George, Emily Goldmann, Stephanie Cook, Fatemeh Choupani, Rubin Baskir, Sydney J. McMaster, Chris Lunt, Karriem Watson, Minnkyong Lee, Sophie Schwartz, Ruchi Munshi, David Glazer, Eric Banks, Anthony Philippakis, Melissa A. Basford, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 28 |
| 2024 | Large language models facilitate the generation of electronic health record phenotyping algorithmsabstractOBJECTIVES: Phenotyping is a core task in observational health research utilizing electronic health records (EHRs). Developing an accurate algorithm demands substantial input from domain experts, involving extensive literature review and evidence synthesis. This burdensome process limits scalability and delays knowledge discovery. We investigate the potential for leveraging large language models (LLMs) to enhance the efficiency of EHR phenotyping by generating high-quality algorithm drafts. MATERIALS AND METHODS: We prompted four LLMs-GPT-4 and GPT-3.5 of ChatGPT, Claude 2, and Bard-in October 2023, asking them to generate executable phenotyping algorithms in the form of SQL queries adhering to a common data model (CDM) for three phenotypes (ie, type 2 diabetes mellitus, dementia, and hypothyroidism). Three phenotyping experts evaluated the returned algorithms across several critical metrics. We further implemented the top-rated algorithms and compared them against clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network. RESULTS: GPT-4 and GPT-3.5 exhibited significantly higher overall expert evaluation scores in instruction following, algorithmic logic, and SQL executability, when compared to Claude 2 and Bard. Although GPT-4 and GPT-3.5 effectively identified relevant clinical concepts, they exhibited immature capability in organizing phenotyping criteria with the proper logic, leading to phenotyping algorithms that were either excessively restrictive (with low recall) or overly broad (with low positive predictive values). CONCLUSION: GPT versions 3.5 and 4 are capable of drafting phenotyping algorithms by identifying relevant clinical criteria aligned with a CDM. However, expertise in informatics and clinical experience is still required to assess and further refine generated algorithms. Chao Yan 0004, Henry H. Ong, Monika E. Grabowska, Matthew S. Krantz, Wu-Chen Su, Alyson L. Dickson, Josh F. Peterson, QiPing Feng, Dan M. Roden, C. Michael Stein, Vern Eric Kerchberger, Bradley A. Malin, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 9 |
| 2024 | PheMIME: an interactive web app and knowledge base for phenome-wide, multi-institutional multimorbidity analysisabstractOBJECTIVES: To address the need for interactive visualization tools and databases in characterizing multimorbidity patterns across different populations, we developed the Phenome-wide Multi-Institutional Multimorbidity Explorer (PheMIME). This tool leverages three large-scale EHR systems to facilitate efficient analysis and visualization of disease multimorbidity, aiming to reveal both robust and novel disease associations that are consistent across different systems and to provide insight for enhancing personalized healthcare strategies. MATERIALS AND METHODS: PheMIME integrates summary statistics from phenome-wide analyses of disease multimorbidities, utilizing data from Vanderbilt University Medical Center, Mass General Brigham, and the UK Biobank. It offers interactive and multifaceted visualizations for exploring multimorbidity. Incorporating an enhanced version of associationSubgraphs, PheMIME also enables dynamic analysis and inference of disease clusters, promoting the discovery of complex multimorbidity patterns. A case study on schizophrenia demonstrates its capability for generating interactive visualizations of multimorbidity networks within and across multiple systems. Additionally, PheMIME supports diverse multimorbidity-based discoveries, detailed further in online case studies. RESULTS: The PheMIME is accessible at https://prod.tbilab.org/PheMIME/. A comprehensive tutorial and multiple case studies for demonstration are available at https://prod.tbilab.org/PheMIME_supplementary_materials/. The source code can be downloaded from https://github.com/tbilab/PheMIME. DISCUSSION: PheMIME represents a significant advancement in medical informatics, offering an efficient solution for accessing, analyzing, and interpreting the complex and noisy real-world patient data in electronic health records. CONCLUSION: PheMIME provides an extensive multimorbidity knowledge base that consolidates data from three EHR systems, and it is a novel interactive tool designed to analyze and visualize multimorbidities across multiple EHR datasets. It stands out as the first of its kind to offer extensive multimorbidity knowledge integration with substantial support for efficient online analysis and interactive visualization. Nick Strayer, Tess Vessels, Karmel Choi, Geoffrey W. Wang, Cosmin Adrian Bejan, Ryan S. Hsi, Alex Bick, Digna R. Velez Edwards, Michael R. Savona, Elizabeth J. Phillips, Jill M. Pulley, Wesley H. Self, Consuelo H. Wilkins, Dan M. Roden, Jordan W. Smoller, Douglas M. Ruderfer, Yaomin Xu |
J. Am. Medical Informatics Assoc. | 16 |
| 2023 | Next-generation phenotyping: introducing phecodeX for enhanced discovery research in medical phenomicsabstractMOTIVATION: Phecodes are widely used and easily adapted phenotypes based on International Classification of Diseases codes. The current version of phecodes (v1.2) was designed primarily to study common/complex diseases diagnosed in adults; however, there are numerous limitations in the codes and their structure. RESULTS: Here, we present phecodeX, an expanded version of phecodes with a revised structure and 1,761 new codes. PhecodeX adds granularity to phenotypes in key disease domains that are under-represented in the current phecode structure-including infectious disease, pregnancy, congenital anomalies, and neonatology-and is a more robust representation of the medical phenome for global use in discovery research. AVAILABILITY AND IMPLEMENTATION: phecodeX is available at https://github.com/PheWAS/phecodeX. Megan M. Shuey, William W. Stead, Ida Aka, April L. Barnado, Lisa Bastarache, Elly Brokamp, Meredith Campbell, Robert J. Carroll, Jeffrey A. Goldstein, Adam Lewis, Beth A. Malow, Jonathan D. Mosley, Travis Osterman, Dolly A Padovani-Claudio, Andrea Ramirez, Dan M. Roden, Bryce A. Schuler, Edward Siew, Jennifer Sucre, Isaac Thomsen, Rory J. Tinker, Sara Van Driest, Colin Walsh, Jeremy L. Warner, Quinn Stanton Wells, Lee E. Wheless |
Bioinform. | 16 |
| 2023 | Interactive network-based clustering and investigation of multimorbidity association matrices with associationSubgraphsabstractMOTIVATION: Making sense of networked multivariate association patterns is vitally important to many areas of high-dimensional analysis. Unfortunately, as the data-space dimensions grow, the number of association pairs increases in O(n2); this means that traditional visualizations such as heatmaps quickly become too complicated to parse effectively. RESULTS: Here, we present associationSubgraphs: a new interactive visualization method to quickly and intuitively explore high-dimensional association datasets using network percolation and clustering. The goal is to provide an efficient investigation of association subgraphs, each containing a subset of variables with stronger and more frequent associations among themselves than the remaining variables outside the subset, by showing the entire clustering dynamics and providing subgraphs under all possible cutoff values at once. Particularly, we apply associationSubgraphs to a phenome-wide multimorbidity association matrix generated from an electronic health record and provide an online, interactive demonstration for exploring multimorbidity subgraphs. AVAILABILITY AND IMPLEMENTATION: An R package implementing both the algorithm and visualization components of associationSubgraphs is available at https://github.com/tbilab/associationsubgraphs. Online documentation is available at https://prod.tbilab.org/associationsubgraphs_info/. A demo using a multimorbidity association matrix is available at https://prod.tbilab.org/associationsubgraphs-example/. Nick Strayer, Lydia Yao, Tess Vessels, Cosmin Adrian Bejan, Ryan S. Hsi, Jana Shirey-Rice, Justin M. Balko, Douglas B. Johnson, Elizabeth J. Phillips, Alex Bick, Todd L. Edwards, Digna R. Velez Edwards, Jill M. Pulley, Quinn Stanton Wells, Michael R. Savona, Nancy J. Cox, Dan M. Roden, Douglas M. Ruderfer, Yaomin Xu |
Bioinform. | 18 |
| 2022 | Comparing medical history data derived from electronic health records and survey answers in the All of Us Research ProgramabstractOBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases. Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | A research agenda to support the development and implementation of genomics-based clinical informatics tools and resourcesabstractOBJECTIVE: The Genomic Medicine Working Group of the National Advisory Council for Human Genome Research virtually hosted its 13th genomic medicine meeting titled "Developing a Clinical Genomic Informatics Research Agenda". The meeting's goal was to articulate a research strategy to develop Genomics-based Clinical Informatics Tools and Resources (GCIT) to improve the detection, treatment, and reporting of genetic disorders in clinical settings. MATERIALS AND METHODS: Experts from government agencies, the private sector, and academia in genomic medicine and clinical informatics were invited to address the meeting's goals. Invitees were also asked to complete a survey to assess important considerations needed to develop a genomic-based clinical informatics research strategy. RESULTS: Outcomes from the meeting included identifying short-term research needs, such as designing and implementing standards-based interfaces between laboratory information systems and electronic health records, as well as long-term projects, such as identifying and addressing barriers related to the establishment and implementation of genomic data exchange systems that, in turn, the research community could help address. DISCUSSION: Discussions centered on identifying gaps and barriers that impede the use of GCIT in genomic medicine. Emergent themes from the meeting included developing an implementation science framework, defining a value proposition for all stakeholders, fostering engagement with patients and partners to develop applications under patient control, promoting the use of relevant clinical workflows in research, and lowering related barriers to regulatory processes. Another key theme was recognizing pervasive biases in data and information systems, algorithms, access, value, and knowledge repositories and identifying ways to resolve them. Ken Wiley, Laura Findley, Madison Goldrich, Teji Rakhra-Burris, Ana Stevens, Pamela Williams, Carol J. Bult, Rex L. Chisholm, Patricia Deverka, Geoffrey S. Ginsburg, Eric D. Green, Gail P. Jarvik, George A. Mensah, Erin Ramos, Mary Relling, Dan M. Roden, Robb Rowley, Gil Alterovitz, Samuel J. Aronson, Lisa Bastarache, James J. Cimino, Erin L. Crowgey, Guilherme Del Fiol, Robert R. Freimuth, Mark A. Hoffman, Janina M. Jeff, Kevin B. Johnson, Kensaku Kawamoto, Subha Madhavan, Eneida A. Mendonça, Lucila Ohno-Machado, Siddharth Pratap, Casey Overby Taylor, Marylyn D. Ritchie, Nephi Walton, Chunhua Weng, Teresa Zayas-Cabán, Teri A. Manolio, Marc S. Williams |
J. Am. Medical Informatics Assoc. | 16 |
| 2021 | Real-time clinical note monitoring to detect conditions for rapid follow-up: A case study of clinical trial enrollment in drug-induced torsades de pointes and Stevens-Johnson syndromeabstractIdentifying acute events as they occur is challenging in large hospital systems. Here, we describe an automated method to detect 2 rare adverse drug events (ADEs), drug-induced torsades de pointes and Stevens-Johnson syndrome and toxic epidermal necrolysis, in near real time for participant recruitment into prospective clinical studies. A text processing system searched clinical notes from the electronic health record (EHR) for relevant keywords and alerted study personnel via email of potential patients for chart review or in-person evaluation. Between 2016 and 2018, the automated recruitment system resulted in capture of 138 true cases of drug-induced rare events, improving recall from 43% to 93%. Our focused electronic alert system maintained 2-year enrollment, including across an EHR migration from a bespoke system to Epic. Real-time monitoring of EHR notes may accelerate research for certain conditions less amenable to conventional study recruitment paradigms. Sarah DeLozier, Peter Speltz, Jason Brito, Leigh Anne Tang, Janey Wang, Joshua C. Smith, Dario A. Giuse, Elizabeth Phillips, Kristina Williams, T. Stephen Strickland, Giovanni Davogustto, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 12 |
| 2021 | DDIWAS: High-throughput electronic health record-based screening of drug-drug interactionsabstractOBJECTIVE: We developed and evaluated Drug-Drug Interaction Wide Association Study (DDIWAS). This novel method detects potential drug-drug interactions (DDIs) by leveraging data from the electronic health record (EHR) allergy list. MATERIALS AND METHODS: To identify potential DDIs, DDIWAS scans for drug pairs that are frequently documented together on the allergy list. Using deidentified medical records, we tested 616 drugs for potential DDIs with simvastatin (a common lipid-lowering drug) and amlodipine (a common blood-pressure lowering drug). We evaluated the performance to rediscover known DDIs using existing knowledge bases and domain expert review. To validate potential novel DDIs, we manually reviewed patient charts and searched the literature. RESULTS: DDIWAS replicated 34 known DDIs. The positive predictive value to detect known DDIs was 0.85 and 0.86 for simvastatin and amlodipine, respectively. DDIWAS also discovered potential novel interactions between simvastatin-hydrochlorothiazide, amlodipine-omeprazole, and amlodipine-valacyclovir. A software package to conduct DDIWAS is publicly available. CONCLUSIONS: In this proof-of-concept study, we demonstrate the value of incorporating information mined from existing allergy lists to detect DDIs in a real-world clinical setting. Since allergy lists are routinely collected in EHRs, DDIWAS has the potential to detect and validate DDI signals across institutions. Patrick Wu, Scott D. Nelson, Juan Zhao 0003, Cosby A. Stone Jr., QiPing Feng, Qingxia Chen, Eric A. Larson, Bingshan Li, Nancy J. Cox, C. Michael Stein, Elizabeth Phillips, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 12 |
| 2021 | Phenotyping coronavirus disease 2019 during a global health pandemic: Lessons learned from the characterization of an early cohort
Sarah DeLozier, Sarah Bland, Melissa McPheeters, Quinn Stanton Wells, Eric Farber-Eger, Cosmin Adrian Bejan, Daniel Fabbri, S. Trent Rosenbloom, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei, Josh F. Peterson, Lisa Bastarache |
J. Biomed. Informatics | 9 |
| 2021 | A retrospective approach to evaluating potential adverse outcomes associated with delay of procedures for cardiovascular and cancer-related diagnoses in the context of COVID-19
Neil S. Zheng, Jeremy L. Warner, Travis Osterman, Quinn Stanton Wells, Xiao-Ou Shu, Steve Deppen, Seth J. Karp, Shon Dwyer, QiPing Feng, Nancy J. Cox, Josh F. Peterson, C. Michael Stein, Dan M. Roden, Kevin B. Johnson, Wei-Qi Wei |
J. Biomed. Informatics | 13 |
| 2020 | Real-time Clinical Note Monitoring to Detect Conditions for Follow-up: a Case Study of Clinical Trial Enrollment in Drug-induced Torsades de Pointes and Stevens-Johnson Syndrome
Sarah DeLozier, Peter Speltz, Jason Brito, Leigh Anne Tang, Janey Wang, Joshua C. Smith, Dario A. Giuse, Elizabeth Phillips, Kristina Williams, Teresa Strickland, Giovanni Davogustto, Dan M. Roden, Joshua C. Denny |
AMIA | 12 |
| 2020 | The All of Us Research Program Researcher Workbench Phenotype Library: Five Disease Implementations
Izabelle P. Humes, Roxana Loperena-Cortes, Melissa A. Basford, Kelsey R. Mayo, Joseph DiPaolo, David J. Schlueter, Wei-Qi Wei, Robert J. Carroll, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden, Andrea H. Ramirez |
AMIA | 13 |
| 2020 | The All of Us Research Program Researcher Workbench: Cloud based access and analytics to advance precision medicine
Andrea H. Ramirez, Kelsey R. Mayo, Robert J. Carroll, Karthik Muthuraman, Melissa A. Basford, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden |
AMIA | 9 |
| 2020 | PheMap: a multi-resource knowledge base for high-throughput phenotyping within electronic health recordsabstractOBJECTIVE: Developing algorithms to extract phenotypes from electronic health records (EHRs) can be challenging and time-consuming. We developed PheMap, a high-throughput phenotyping approach that leverages multiple independent, online resources to streamline the phenotyping process within EHRs. MATERIALS AND METHODS: PheMap is a knowledge base of medical concepts with quantified relationships to phenotypes that have been extracted by natural language processing from publicly available resources. PheMap searches EHRs for each phenotype's quantified concepts and uses them to calculate an individual's probability of having this phenotype. We compared PheMap to clinician-validated phenotyping algorithms from the Electronic Medical Records and Genomics (eMERGE) network for type 2 diabetes mellitus (T2DM), dementia, and hypothyroidism using 84 821 individuals from Vanderbilt Univeresity Medical Center's BioVU DNA Biobank. We implemented PheMap-based phenotypes for genome-wide association studies (GWAS) for T2DM, dementia, and hypothyroidism, and phenome-wide association studies (PheWAS) for variants in FTO, HLA-DRB1, and TCF7L2. RESULTS: In this initial iteration, the PheMap knowledge base contains quantified concepts for 841 disease phenotypes. For T2DM, dementia, and hypothyroidism, the accuracy of the PheMap phenotypes were >97% using a 50% threshold and eMERGE case-control status as a reference standard. In the GWAS analyses, PheMap-derived phenotype probabilities replicated 43 of 51 previously reported disease-associated variants for the 3 phenotypes. For 9 of the 11 top associations, PheMap provided an equivalent or more significant P value than eMERGE-based phenotypes. The PheMap-based PheWAS showed comparable or better performance to a traditional phecode-based PheWAS. PheMap is publicly available online. CONCLUSIONS: PheMap significantly streamlines the process of extracting research-quality phenotype information from EHRs, with comparable or better performance to current phenotyping approaches. Neil S. Zheng, QiPing Feng, Vern Eric Kerchberger, Juan Zhao 0003, Todd L. Edwards, Nancy J. Cox, C. Michael Stein, Dan M. Roden, Joshua C. Denny, Wei-Qi Wei |
J. Am. Medical Informatics Assoc. | 8 |
| 2019 | Improving the phenotype risk score as a scalable approach to identifying patients with Mendelian diseaseabstractOBJECTIVE: The Phenotype Risk Score (PheRS) is a method to detect Mendelian disease patterns using phenotypes from the electronic health record (EHR). We compared the performance of different approaches mapping EHR phenotypes to Mendelian disease features. MATERIALS AND METHODS: PheRS utilizes Mendelian diseases descriptions annotated with Human Phenotype Ontology (HPO) terms. In previous work, we presented a map linking phecodes (based on International Classification of Diseases [ICD]-Ninth Revision) to HPO terms. For this study, we integrated ICD-Tenth Revision codes and lab data. We also created a new map between HPO terms using customized groupings of ICD codes. We compared the performance with cases and controls for 16 Mendelian diseases using 2.5 million de-identified medical records. RESULTS: PheRS effectively distinguished cases from controls for all 15 positive controls and all approaches tested (P < 4 × 1016). Adding lab data led to a statistically significant improvement for 4 of 14 diseases. The custom ICD groupings improved specificity, leading to an average 8% increase for precision at 100 (-2% to 22%). Eight of 10 adults with cystic fibrosis tested had PheRS in the 95th percentile prio to diagnosis. DISCUSSION: Both phecodes and custom ICD groupings were able to detect differences between affected cases and controls at the population level. The ICD map showed better precision for the highest scoring individuals. Adding lab data improved performance at detecting population-level differences. CONCLUSIONS: PheRS is a scalable method to study Mendelian disease at the population level using electronic health record data and can potentially be used to find patients with undiagnosed Mendelian disease. Lisa Bastarache, Jacob J. Hughey, Jeffery A. Goldstein, Julie A. Bastraache, Satya Das, Neil Charles Zaki, Chenjie Zeng, Leigh Anne Tang, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 9 |
| 2018 | Evaluating statistical approaches to leverage large clinical datasets for uncovering therapeutic and adverse medication effectsabstractMotivation: Phenome-wide association studies (PheWAS) have been used to discover many genotype-phenotype relationships and have the potential to identify therapeutic and adverse drug outcomes using longitudinal data within electronic health records (EHRs). However, the statistical methods for PheWAS applied to longitudinal EHR medication data have not been established. Results: In this study, we developed methods to address two challenges faced with reuse of EHR for this purpose: confounding by indication, and low exposure and event rates. We used Monte Carlo simulation to assess propensity score (PS) methods, focusing on two of the most commonly used methods, PS matching and PS adjustment, to address confounding by indication. We also compared two logistic regression approaches (the default of Wald versus Firth's penalized maximum likelihood, PML) to address complete separation due to sparse data with low exposure and event rates. PS adjustment resulted in greater power than PS matching, while controlling Type I error at 0.05. The PML method provided reasonable P-values, even in cases with complete separation, with well controlled Type I error rates. Using PS adjustment and the PML method, we identify novel latent drug effects in pediatric patients exposed to two common antibiotic drugs, ampicillin and gentamicin. Availability and implementation: R packages PheWAS and EHR are available at https://github.com/PheWAS/PheWAS and at CRAN (https://www.r-project.org/), respectively. The R script for data processing and the main analysis is available at https://github.com/choileena/EHR. Supplementary information: Supplementary data are available at Bioinformatics online. Leena Choi, Robert J. Carroll, Cole Beck, Jonathan D. Mosley, Dan M. Roden, Joshua C. Denny, Sara L. Van Driest |
Bioinform. | 5 |
| 2018 | A case study evaluating the portability of an executable computable phenotype algorithm across multiple institutions and electronic health record environmentsabstractElectronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges. Jennifer A. Pacheco, Luke V. Rasmussen, Richard C. Kiefer, Thomas R. Campion Jr., Peter Speltz, Robert J. Carroll, Sarah C. Stallings, Huan Mo, Monika Ahuja, Guoqian Jiang, Eric LaRose, Peggy L. Peissig, Ning Shang 0004, Barbara Benoit, Vivian S. Gainer, Kenneth Borthwick, Kathryn L. Jackson, Ambrish Sharma, Andy Yizhou Wu, Abel N. Kho, Dan M. Roden, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
J. Am. Medical Informatics Assoc. | 21 |
| 2017 | Evaluating electronic health record data sources and algorithmic approaches to identify hypertensive individualsabstractOBJECTIVE: Phenotyping algorithms applied to electronic health record (EHR) data enable investigators to identify large cohorts for clinical and genomic research. Algorithm development is often iterative, depends on fallible investigator intuition, and is time- and labor-intensive. We developed and evaluated 4 types of phenotyping algorithms and categories of EHR information to identify hypertensive individuals and controls and provide a portable module for implementation at other sites. MATERIALS AND METHODS: We reviewed the EHRs of 631 individuals followed at Vanderbilt for hypertension status. We developed features and phenotyping algorithms of increasing complexity. Input categories included International Classification of Diseases, Ninth Revision (ICD9) codes, medications, vital signs, narrative-text search results, and Unified Medical Language System (UMLS) concepts extracted using natural language processing (NLP). We developed a module and tested portability by replicating 10 of the best-performing algorithms at the Marshfield Clinic. RESULTS: Random forests using billing codes, medications, vitals, and concepts had the best performance with a median area under the receiver operator characteristic curve (AUC) of 0.976. Normalized sums of all 4 categories also performed well (0.959 AUC). The best non-NLP algorithm combined normalized ICD9 codes, medications, and blood pressure readings with a median AUC of 0.948. Blood pressure cutoffs or ICD9 code counts alone had AUCs of 0.854 and 0.908, respectively. Marshfield Clinic results were similar. CONCLUSION: This work shows that billing codes or blood pressure readings alone yield good hypertension classification performance. However, even simple combinations of input categories improve performance. The most complex algorithms classified hypertension with excellent recall and precision. Pedro L. Teixeira, Wei-Qi Wei, Robert M. Cronin, Huan Mo, Jacob P. VanHouten, Robert J. Carroll, Eric LaRose, Lisa Bastarache, S. Trent Rosenbloom, Todd L. Edwards, Dan M. Roden, Thomas A. Lasko, Richard A. Dart, Anne M. Nikolai, Peggy L. Peissig, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 11 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 16 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 28 |
| 2015 | Validating drug repurposing signals using electronic health records: a case study of metformin associated with reduced cancer mortalityabstractOBJECTIVES: Drug repurposing, which finds new indications for existing drugs, has received great attention recently. The goal of our work is to assess the feasibility of using electronic health records (EHRs) and automated informatics methods to efficiently validate a recent drug repurposing association of metformin with reduced cancer mortality. METHODS: By linking two large EHRs from Vanderbilt University Medical Center and Mayo Clinic to their tumor registries, we constructed a cohort including 32,415 adults with a cancer diagnosis at Vanderbilt and 79,258 cancer patients at Mayo from 1995 to 2010. Using automated informatics methods, we further identified type 2 diabetes patients within the cancer cohort and determined their drug exposure information, as well as other covariates such as smoking status. We then estimated HRs for all-cause mortality and their associated 95% CIs using stratified Cox proportional hazard models. HRs were estimated according to metformin exposure, adjusted for age at diagnosis, sex, race, body mass index, tobacco use, insulin use, cancer type, and non-cancer Charlson comorbidity index. RESULTS: Among all Vanderbilt cancer patients, metformin was associated with a 22% decrease in overall mortality compared to other oral hypoglycemic medications (HR 0.78; 95% CI 0.69 to 0.88) and with a 39% decrease compared to type 2 diabetes patients on insulin only (HR 0.61; 95% CI 0.50 to 0.73). Diabetic patients on metformin also had a 23% improved survival compared with non-diabetic patients (HR 0.77; 95% CI 0.71 to 0.85). These associations were replicated using the Mayo Clinic EHR data. Many site-specific cancers including breast, colorectal, lung, and prostate demonstrated reduced mortality with metformin use in at least one EHR. CONCLUSIONS: EHR data suggested that the use of metformin was associated with decreased mortality after a cancer diagnosis compared with diabetic and non-diabetic cancer patients not on metformin, indicating its potential as a chemotherapeutic regimen. This study serves as a model for robust and inexpensive validation studies for drug repurposing signals using EHR data. Hua Xu 0001, Melinda Aldrich, Qingxia Chen, Neeraja B. Peterson, Mia A. Levy, Anushi Shah, Xiaoyang Ruan, Min Jiang 0007, Jamii St Julien, Jeremy L. Warner, Carol Friedman, Dan M. Roden, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 16 |
| 2015 | Deciphering Signaling Pathway Networks to Understand the Molecular Mechanisms of Metformin ActionabstractA drug exerts its effects typically through a signal transduction cascade, which is non-linear and involves intertwined networks of multiple signaling pathways. Construction of such a signaling pathway network (SPNetwork) can enable identification of novel drug targets and deep understanding of drug action. However, it is challenging to synopsize critical components of these interwoven pathways into one network. To tackle this issue, we developed a novel computational framework, the Drug-specific Signaling Pathway Network (DSPathNet). The DSPathNet amalgamates the prior drug knowledge and drug-induced gene expression via random walk algorithms. Using the drug metformin, we illustrated this framework and obtained one metformin-specific SPNetwork containing 477 nodes and 1,366 edges. To evaluate this network, we performed the gene set enrichment analysis using the disease genes of type 2 diabetes (T2D) and cancer, one T2D genome-wide association study (GWAS) dataset, three cancer GWAS datasets, and one GWAS dataset of cancer patients with T2D on metformin. The results showed that the metformin network was significantly enriched with disease genes for both T2D and cancer, and that the network also included genes that may be associated with metformin-associated cancer survival. Furthermore, from the metformin SPNetwork and common genes to T2D and cancer, we generated a subnetwork to highlight the molecule crosstalk between T2D and cancer. The follow-up network analyses and literature mining revealed that seven genes (CDKN1A, ESR1, MAX, MYC, PPARGC1A, SP1, and STK11) and one novel MYC-centered pathway with CDKN1A, SP1, and STK11 might play important roles in metformin's antidiabetic and anticancer effects. Some results are supported by previous studies. In summary, our study 1) develops a novel framework to construct drug-specific signal transduction networks; 2) provides insights into the molecular mode of metformin; 3) serves a model for exploring signaling pathways to facilitate understanding of drug action, disease pathogenesis, and identification of drug targets. Jingchun Sun, Min Zhao 0006, Peilin Jia, Lily Wang 0001, Yonghui Wu 0001, Carissa Iverson, Yubo Zhou, Erica A. Bowton, Dan M. Roden, Joshua C. Denny, Melinda Aldrich, Hua Xu 0001, Zhongming Zhao |
PLoS Comput. Biol. | 9 |
| 2014 | Replication of SCN5A Associations with Electrocardiographic Traits in African Americans from Clinical and Epidemiologic Studies
Janina M. Jeff, Kristin Brown-Gentry, Robert J. Goodloe, Marylyn D. Ritchie, Joshua C. Denny, Abel N. Kho, Loren L. Armstrong, Bob McClellan Jr., Ping Mayo, Hailing Jin, Niloufar B. Gillani, Nathalie Schnetz-Boutaud, Holli H. Dilks, Melissa A. Basford, Jennifer A. Pacheco, Gail P. Jarvik, Rex L. Chisholm, Dan M. Roden, M. Geoffrey Hayes, Dana C. Crawford |
EvoApplications | 19 |
| 2014 | Size matters: How population size influences genotype-phenotype association studies in anonymized data
Raymond Heatherly, Joshua C. Denny, Jonathan L. Haines, Dan M. Roden, Bradley A. Malin |
J. Biomed. Informatics | 4 |
| 2011 | Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarinabstractOBJECTIVE: DNA biobanks linked to comprehensive electronic health records systems are potentially powerful resources for pharmacogenetic studies. This study sought to develop natural-language-processing algorithms to extract drug-dose information from clinical text, and to assess the capabilities of such tools to automate the data-extraction process for pharmacogenetic studies. MATERIALS AND METHODS: A manually validated warfarin pharmacogenetic study identified a cohort of 1125 patients with a stable warfarin dose, in which 776 patients were managed by Coumadin Clinic physicians, and the remaining 349 patients were managed by their providers. The authors developed two algorithms to extract weekly warfarin doses from both data sets: a regular expression-based program for semistructured Coumadin Clinic notes; and an advanced weekly dose calculator based on an existing medication information extraction system (MedEx) for narrative providers' notes. The authors then conducted an association analysis between an automatically extracted stable weekly dose of warfarin and four genetic variants of VKORC1 and CYP2C9 genes. The performance of the weekly dose-extraction program was evaluated by comparing it with a gold standard containing manually curated weekly doses. Precision, recall, F-measure, and overall accuracy were reported. Associations between known variants in VKORC1 and CYP2C9 and warfarin stable weekly dose were performed with linear regression adjusted for age, gender, and body mass index. RESULTS: The authors' evaluation showed that the MedEx-based system could determine patients' warfarin weekly doses with 99.7% recall, 90.8% precision, and 93.8% accuracy. Using the automatically extracted weekly doses of warfarin, the authors successfully replicated the previous known associations between warfarin stable dose and genetic variants in VKORC1 and CYP2C9. Hua Xu 0001, Min Jiang 0007, Matthew Oetjens, Erica A. Bowton, Andrea H. Ramirez, Janina M. Jeff, Melissa A. Basford, Jill M. Pulley, James D. Cowan, Marylyn D. Ritchie, Daniel R. Masys, Dan M. Roden, Dana C. Crawford, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 13 |
| 2010 | PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene-disease associationsabstractMOTIVATION: Emergence of genetic data coupled to longitudinal electronic medical records (EMRs) offers the possibility of phenome-wide association scans (PheWAS) for disease-gene associations. We propose a novel method to scan phenomic data for genetic associations using International Classification of Disease (ICD9) billing codes, which are available in most EMR systems. We have developed a code translation table to automatically define 776 different disease populations and their controls using prevalent ICD9 codes derived from EMR data. As a proof of concept of this algorithm, we genotyped the first 6005 European-Americans accrued into BioVU, Vanderbilt's DNA biobank, at five single nucleotide polymorphisms (SNPs) with previously reported disease associations: atrial fibrillation, Crohn's disease, carotid artery stenosis, coronary artery disease, multiple sclerosis, systemic lupus erythematosus and rheumatoid arthritis. The PheWAS software generated cases and control populations across all ICD9 code groups for each of these five SNPs, and disease-SNP associations were analyzed. The primary outcome of this study was replication of seven previously known SNP-disease associations for these SNPs. RESULTS: Four of seven known SNP-disease associations using the PheWAS algorithm were replicated with P-values between 2.8 x 10(-6) and 0.011. The PheWAS algorithm also identified 19 previously unknown statistical associations between these SNPs and diseases at P < 0.01. This study indicates that PheWAS analysis is a feasible method to investigate SNP-disease associations. Further evaluation is needed to determine the validity of these associations and the appropriate statistical thresholds for clinical significance. AVAILABILITY: The PheWAS software and code translation table are freely available at http://knowledgemap.mc.vanderbilt.edu/research. Joshua C. Denny, Marylyn D. Ritchie, Melissa A. Basford, Jill M. Pulley, Lisa Bastarache, Kristin Brown-Gentry, Deede Wang, Daniel R. Masys, Dan M. Roden, Dana C. Crawford |
Bioinform. | 9 |
| 2010 | An analytical approach to characterize morbidity profile dissimilarity between distinct cohorts using electronic medical records
Jonathan S. Schildcrout, Melissa A. Basford, Jill M. Pulley, Daniel R. Masys, Dan M. Roden, Deede Wang, Christopher G. Chute, Iftikhar J. Kullo, David Carrell, Peggy L. Peissig, Abel N. Kho, Joshua C. Denny |
J. Biomed. Informatics | 5 |