Robert J. Carroll

dblp:119/7440 · DBLP profile ↗
← Back
36ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0003-3802-8183ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 7 first-author · 13 since 2021
YearPublicationVenuePosition
2025 PheWAS analysis on large-scale biobank data with PheTK
abstract
SUMMARY: With the rapid growth of genetic data linked to electronic health record (EHR) data in huge cohorts, large-scale phenome-wide association study (PheWAS) have become powerful discovery tools in biomedical research. PheWAS is an analysis method to study phenotype associations utilizing longitudinal EHR data. Previous PheWAS packages were developed mostly with smaller datasets and with earlier PheWAS approaches. PheTK was designed to simplify analysis and efficiently handle biobank-scale data. PheTK uses multithreading and supports a full PheWAS workflow including extraction of data from OMOP databases and Hail matrix tables as well as PheWAS analysis for both phecode version 1.2 and phecodeX. Benchmarking results showed PheTK took 64% less time than the R PheWAS package to complete the same workflow. PheTK can be run locally or on cloud platforms such as the All of Us Researcher Workbench (All of Us) or the UK Biobank (UKB) Research Analysis Platform (RAP). AVAILABILITY AND IMPLEMENTATION: The PheTK package is freely available on the Python Package Index, on GitHub under GNU General Public License (GPL-3) at https://github.com/nhgritctran/PheTK, and on Zenodo, DOI 10.5281/zenodo.14217954, at https://doi.org/10.5281/zenodo.14217954. PheTK is implemented in Python and platform independent.
Tam C. Tran, David J. Schlueter, Chenjie Zeng, Huan Mo, Robert J. Carroll, Joshua C. Denny
Bioinform.5
2024 Developing and evaluating pediatric phecodes (Peds-Phecodes) for high-throughput phenotyping using electronic health records
abstract
OBJECTIVE: Pediatric patients have different diseases and outcomes than adults; however, existing phecodes do not capture the distinctive pediatric spectrum of disease. We aim to develop specialized pediatric phecodes (Peds-Phecodes) to enable efficient, large-scale phenotypic analyses of pediatric patients. MATERIALS AND METHODS: We adopted a hybrid data- and knowledge-driven approach leveraging electronic health records (EHRs) and genetic data from Vanderbilt University Medical Center to modify the most recent version of phecodes to better capture pediatric phenotypes. First, we compared the prevalence of patient diagnoses in pediatric and adult populations to identify disease phenotypes differentially affecting children and adults. We then used clinical domain knowledge to remove phecodes representing phenotypes unlikely to affect pediatric patients and create new phecodes for phenotypes relevant to the pediatric population. We further compared phenome-wide association study (PheWAS) outcomes replicating known pediatric genotype-phenotype associations between Peds-Phecodes and phecodes. RESULTS: The Peds-Phecodes aggregate 15 533 ICD-9-CM codes and 82 949 ICD-10-CM codes into 2051 distinct phecodes. Peds-Phecodes replicated more known pediatric genotype-phenotype associations than phecodes (248 vs 192 out of 687 SNPs, P < .001). DISCUSSION: We introduce Peds-Phecodes, a high-throughput EHR phenotyping tool tailored for use in pediatric populations. We successfully validated the Peds-Phecodes using genetic replication studies. Our findings also reveal the potential use of Peds-Phecodes in detecting novel genotype-phenotype associations for pediatric conditions. We expect that Peds-Phecodes will facilitate large-scale phenomic and genomic analyses in pediatric populations. CONCLUSION: Peds-Phecodes capture higher-quality pediatric phenotypes and deliver superior PheWAS outcomes compared to phecodes.
Monika E. Grabowska, Sara L. Van Driest, Jamie R. Robinson, Anna E. Patrick, Chris Guardo, Srushti Gangireddy, Henry H. Ong, QiPing Feng, Robert J. Carroll, Prince J. Kannankeril, Wei-Qi Wei
J. Am. Medical Informatics Assoc.9
2024 Identifying erroneous height and weight values from adult electronic health records in the All of Us research program
abstract
INTRODUCTION: Electronic Health Records (EHR) are a useful data source for research, but their usability is hindered by measurement errors. This study investigated an automatic error detection algorithm for adult height and weight measurements in EHR for the All of Us Research Program (All of Us). METHODS: We developed reference charts for adult heights and weights that were stratified on participant sex. Our analysis included 4,076,534 height and 5,207,328 wt measurements from ∼ 150,000 participants. Errors were identified using modified standard deviation scores, differences from their expected values, and significant changes between consecutive measurements. We evaluated our method with chart-reviewed heights (8,092) and weights (9,039) from 250 randomly selected participants and compared it with the current cleaning algorithm in All of Us. RESULTS: The proposed algorithm classified 1.4 % of height and 1.5 % of weight errors in the full cohort. Sensitivity was 90.4 % (95 % CI: 79.0-96.8 %) for heights and 65.9 % (95 % CI: 56.9-74.1 %) for weights. Precision was 73.4 % (95 % CI: 60.9-83.7 %) for heights and 62.9 (95 % CI: 54.0-71.1 %) for weights. In comparison, the current cleaning algorithm has inferior performance in sensitivity (55.8 %) and precision (16.5 %) for height errors while having higher precision (94.0 %) and lower sensitivity (61.9 %) for weight errors. DISCUSSION: Our proposed algorithm outperformed in detecting height errors compared to weights. It can serve as a valuable addition to the current All of Us cleaning algorithm for identifying erroneous height values.
Andrew Guide, Lina M. Sulieman, Shawn Garbett, Robert M. Cronin, Matthew E. Spotnitz, Karthik Natarajan, Robert J. Carroll, Paul A. Harris, Qingxia Chen
J. Biomed. Informatics7
2023 Next-generation phenotyping: introducing phecodeX for enhanced discovery research in medical phenomics
abstract
MOTIVATION: Phecodes are widely used and easily adapted phenotypes based on International Classification of Diseases codes. The current version of phecodes (v1.2) was designed primarily to study common/complex diseases diagnosed in adults; however, there are numerous limitations in the codes and their structure. RESULTS: Here, we present phecodeX, an expanded version of phecodes with a revised structure and 1,761 new codes. PhecodeX adds granularity to phenotypes in key disease domains that are under-represented in the current phecode structure-including infectious disease, pregnancy, congenital anomalies, and neonatology-and is a more robust representation of the medical phenome for global use in discovery research. AVAILABILITY AND IMPLEMENTATION: phecodeX is available at https://github.com/PheWAS/phecodeX.
Megan M. Shuey, William W. Stead, Ida Aka, April L. Barnado, Lisa Bastarache, Elly Brokamp, Meredith Campbell, Robert J. Carroll, Jeffrey A. Goldstein, Adam Lewis, Beth A. Malow, Jonathan D. Mosley, Travis Osterman, Dolly A Padovani-Claudio, Andrea Ramirez, Dan M. Roden, Bryce A. Schuler, Edward Siew, Jennifer Sucre, Isaac Thomsen, Rory J. Tinker, Sara Van Driest, Colin Walsh, Jeremy L. Warner, Quinn Stanton Wells, Lee E. Wheless
Bioinform.8
2023 Characterizing variability of electronic health record-driven phenotype definitions
abstract
OBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic.
Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen
J. Am. Medical Informatics Assoc.17
2023 Managing re-identification risks while providing access to the All of Us research program
abstract
OBJECTIVE: The All of Us Research Program makes individual-level data available to researchers while protecting the participants' privacy. This article describes the protections embedded in the multistep access process, with a particular focus on how the data was transformed to meet generally accepted re-identification risk levels. METHODS: At the time of the study, the resource consisted of 329 084 participants. Systematic amendments were applied to the data to mitigate re-identification risk (eg, generalization of geographic regions, suppression of public events, and randomization of dates). We computed the re-identification risk for each participant using a state-of-the-art adversarial model specifically assuming that it is known that someone is a participant in the program. We confirmed the expected risk is no greater than 0.09, a threshold that is consistent with guidelines from various US state and federal agencies. We further investigated how risk varied as a function of participant demographics. RESULTS: The results indicated that 95th percentile of the re-identification risk of all the participants is below current thresholds. At the same time, we observed that risk levels were higher for certain race, ethnic, and genders. CONCLUSIONS: While the re-identification risk was sufficiently low, this does not imply that the system is devoid of risk. Rather, All of Us uses a multipronged data protection strategy that includes strong authentication practices, active monitoring of data misuse, and penalization mechanisms for users who violate terms of service.
Weiyi Xia, Melissa A. Basford, Robert J. Carroll, Ellen Wright Clayton, Paul A. Harris, Murat Kantarcioglu, Yongtai Liu, Steve Nyemba, Yevgeniy Vorobeychik, Zhiyu Wan, Bradley A. Malin
J. Am. Medical Informatics Assoc.3
2022 Self-paced Training Modality to Promote the Use of All of Us Researcher Workbench in Educational and Research Settings
Hiral Master, Lina M. Sulieman, Paul A. Harris, Karthik Natarajan, Robert J. Carroll, Kayla Marginean, Kelsey R. Mayo, Aymone Kouame
AMIA5
2022 Assessing Data Quality and Diversity in the All of Us
Lina M. Sulieman, Jennifer Zhang, Kayla Marginean, Paul A. Harris, Robert J. Carroll
AMIA5
2022 Harmonizing FHIR and Common Data Models Used in Research: Current State and a Path Forward
Teresa Zayas-Cabán, Belinda Seto, Robert J. Carroll, Jon Duke, Viet Nguyen
AMIA3
2022 Inference-based correction of multi-site height and weight measurement data in the All of Us research program
abstract
OBJECTIVE: Measurement and data entry of height and weight values are error prone. Aggregation of medical record data from multiple sites creates new challenges prompting the need to identify and correct errant values. We sought to characterize and correct issues with height and weight measurement values within the All of Us (AoU) Research Program. MATERIALS AND METHODS: Using the AoU Researcher Workbench, we assessed site-level measurement value distributions to infer unit types. We also used plausibility checks with exceptions for conditions with possible outlier values, eg obesity, and assessed for excess deviation within individual participant's records. RESULTS: 15.8% of height and 22.4% of weight values had missing unit type information. DISCUSSION: We identified several measurement unit related issues: the use of different units of measure within and between sites, missing units, and incorrect labeling of units. Failure to account for these in patient data repositories may lead to erroneous study results and conclusions. CONCLUSION: Discrepancies in height and weight measurement data may arise from missing or mislabeled units. Using site- and participant-level analyses while accounting for outlier value-associated clinical conditions, we can infer measurement units and apply corrections. These methods are adaptable and expandable within AoU and other data repositories.
Mirza S. Khan, Robert J. Carroll
J. Am. Medical Informatics Assoc.2
2022 Comparing medical history data derived from electronic health records and survey answers in the All of Us Research Program
abstract
OBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases.
Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez
J. Am. Medical Informatics Assoc.3
2021 Measuring the correctness of All of Us physical measurement
Lina M. Sulieman, Karthik Natarajan, Qingxia Chen, Robert J. Carroll, Kayla Marginean, Paul A. Harris, Andrea H. Ramirez
AMIA4
2021 Comparison of family health history in surveys vs electronic health record data mapped to the observational medical outcomes partnership data model in the All of Us Research Program
abstract
OBJECTIVE: Family health history is important to clinical care and precision medicine. Prior studies show gaps in data collected from patient surveys and electronic health records (EHRs). The All of Us Research Program collects family history from participants via surveys and EHRs. This Demonstration Project aims to evaluate availability of family health history information within the publicly available data from All of Us and to characterize the data from both sources. MATERIALS AND METHODS: Surveys were completed by participants on an electronic portal. EHR data was mapped to the Observational Medical Outcomes Partnership data model. We used descriptive statistics to perform exploratory analysis of the data, including evaluating a list of medically actionable genetic disorders. We performed a subanalysis on participants who had both survey and EHR data. RESULTS: There were 54 872 participants with family history data. Of those, 26% had EHR data only, 63% had survey only, and 10.5% had data from both sources. There were 35 217 participants with reported family history of a medically actionable genetic disorder (9% from EHR only, 89% from surveys, and 2% from both). In the subanalysis, we found inconsistencies between the surveys and EHRs. More details came from surveys. When both mentioned a similar disease, the source of truth was unclear. CONCLUSIONS: Compiling data from both surveys and EHR can provide a more comprehensive source for family health history, but informatics challenges and opportunities exist. Access to more complete understanding of a person's family health history may provide opportunities for precision medicine.
Robert M. Cronin, Alese E. Halvorson, Cassie Springer, Xiaoke Feng, Lina M. Sulieman, Roxana Loperena-Cortes, Kelsey R. Mayo, Robert J. Carroll, Qingxia Chen, Brian K. Ahmedani, Jason Karnes, Bruce Korf, Christopher J. O'Donnell, Andrea H. Ramirez
J. Am. Medical Informatics Assoc.8
2020 Curating Data and Communicating Quality for Impact in a FAIR World
Robert J. Carroll, Kristin Wuichet, Charles Phillips, Michelle Holko, Allison P. Heath
AMIA1
2020 The All of Us Research Program Researcher Workbench Phenotype Library: Five Disease Implementations
Izabelle P. Humes, Roxana Loperena-Cortes, Melissa A. Basford, Kelsey R. Mayo, Joseph DiPaolo, David J. Schlueter, Wei-Qi Wei, Robert J. Carroll, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden, Andrea H. Ramirez
AMIA9
2020 The All of Us Research Program Researcher Workbench: Cloud based access and analytics to advance precision medicine
Andrea H. Ramirez, Kelsey R. Mayo, Robert J. Carroll, Karthik Muthuraman, Melissa A. Basford, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden
AMIA3
2019 Curating EHR data in the All of Us Research Program
Karthik Natarajan, Robert J. Carroll, Thomas R. Campion Jr., Joan Grand, Shyam Visweswaran
AMIA2
2019 Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng
J. Biomed. Informatics7
2019 Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng
J. Biomed. Informatics5
2019 Automated grouping of medical codes via multiview banded spectral clustering
abstract
OBJECTIVE: With its increasingly widespread adoption, electronic health records (EHR) have enabled phenotypic information extraction at an unprecedented granularity and scale. However, often a medical concept (e.g. diagnosis, prescription, symptom) is described in various synonyms across different EHR systems, hindering data integration for signal enhancement and complicating dimensionality reduction for knowledge discovery. Despite existing ontologies and hierarchies, tremendous human effort is needed for curation and maintenance - a process that is both unscalable and susceptible to subjective biases. This paper aims to develop a data-driven approach to automate grouping medical terms into clinically relevant concepts by combining multiple up-to-date data sources in an unbiased manner. METHODS: We present a novel data-driven grouping approach - multi-view banded spectral clustering (mvBSC) combining summary data from multiple healthcare systems. The proposed method consists of a banding step that leverages the prior knowledge from the existing coding hierarchy, and a combining step that performs spectral clustering on an optimally weighted matrix. RESULTS: -measure, and were found to consistently exhibit great similarity to the existing manual grouping counterpart. The resulting ICD groupings also enjoy comparable interpretability and are well aligned with the current ICD hierarchy. CONCLUSION: The proposed approach, by systematically leveraging multiple data sources, is able to overcome bias while maximizing consensus to achieve generalizability. It has the advantage of being efficient, scalable, and adaptive to the evolving human knowledge reflected in the data, showing a significant step toward automating medical knowledge integration.
Luwan Zhang, Tianrun A. Cai, Yuri Ahuja, Zeling He, Yuk-Lam Ho, Andrew L. Beam, Kelly Cho, Robert J. Carroll, Joshua C. Denny, Isaac S. Kohane, Katherine P. Liao, Tianxi Cai
J. Biomed. Informatics9
2018 Evaluating statistical approaches to leverage large clinical datasets for uncovering therapeutic and adverse medication effects
abstract
Motivation: Phenome-wide association studies (PheWAS) have been used to discover many genotype-phenotype relationships and have the potential to identify therapeutic and adverse drug outcomes using longitudinal data within electronic health records (EHRs). However, the statistical methods for PheWAS applied to longitudinal EHR medication data have not been established. Results: In this study, we developed methods to address two challenges faced with reuse of EHR for this purpose: confounding by indication, and low exposure and event rates. We used Monte Carlo simulation to assess propensity score (PS) methods, focusing on two of the most commonly used methods, PS matching and PS adjustment, to address confounding by indication. We also compared two logistic regression approaches (the default of Wald versus Firth's penalized maximum likelihood, PML) to address complete separation due to sparse data with low exposure and event rates. PS adjustment resulted in greater power than PS matching, while controlling Type I error at 0.05. The PML method provided reasonable P-values, even in cases with complete separation, with well controlled Type I error rates. Using PS adjustment and the PML method, we identify novel latent drug effects in pediatric patients exposed to two common antibiotic drugs, ampicillin and gentamicin. Availability and implementation: R packages PheWAS and EHR are available at https://github.com/PheWAS/PheWAS and at CRAN (https://www.r-project.org/), respectively. The R script for data processing and the main analysis is available at https://github.com/choileena/EHR. Supplementary information: Supplementary data are available at Bioinformatics online.
Leena Choi, Robert J. Carroll, Cole Beck, Jonathan D. Mosley, Dan M. Roden, Joshua C. Denny, Sara L. Van Driest
Bioinform.2
2018 Uncovering exposures responsible for birth season - disease effects: a global study
abstract
OBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them.
Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti
J. Am. Medical Informatics Assoc.5
2018 A case study evaluating the portability of an executable computable phenotype algorithm across multiple institutions and electronic health record environments
abstract
Electronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges.
Jennifer A. Pacheco, Luke V. Rasmussen, Richard C. Kiefer, Thomas R. Campion Jr., Peter Speltz, Robert J. Carroll, Sarah C. Stallings, Huan Mo, Monika Ahuja, Guoqian Jiang, Eric LaRose, Peggy L. Peissig, Ning Shang 0004, Barbara Benoit, Vivian S. Gainer, Kenneth Borthwick, Kathryn L. Jackson, Ambrish Sharma, Andy Yizhou Wu, Abel N. Kho, Dan M. Roden, Jyotishman Pathak, Joshua C. Denny, William K. Thompson
J. Am. Medical Informatics Assoc.6
2017 The Data and Research Center of the All of Us Research Program: Framework for a National Cohort Program and Research Opportunities
Robert J. Carroll, Joshua C. Mandel, Karthik Natarajan, Scott Sutherland, Joshua C. Denny
AMIA1
2017 HealthPro: An integrated web application for essential health data and biological specimen collection in the Precision Medicine Initiative
Kelsey R. Mayo, Robert J. Carroll, Jason Tan, Rebecca Johnston, Celecia M. Scott, Joshua C. Denny, Paul A. Harris
AMIA2
2017 Sub-Phenotyping of Crohn's Disease Using a Large Electronic Record Cohort
Jamie R. Robinson, Lisa Bastarache, Robert J. Carroll, Elizabeth A. Scoville, David A. Schwartz, Joshua C. Denny
AMIA3
2017 Association of BMI and Obesity Genetic Risk Score with Surgical Procedures Through a Procedure-wide Association Study
Jamie R. Robinson, Zongyang Mou, Lisa Bastarache, Wei-Qi Wei, Robert J. Carroll, Joshua C. Denny
AMIA5
2017 Evaluating electronic health record data sources and algorithmic approaches to identify hypertensive individuals
abstract
OBJECTIVE: Phenotyping algorithms applied to electronic health record (EHR) data enable investigators to identify large cohorts for clinical and genomic research. Algorithm development is often iterative, depends on fallible investigator intuition, and is time- and labor-intensive. We developed and evaluated 4 types of phenotyping algorithms and categories of EHR information to identify hypertensive individuals and controls and provide a portable module for implementation at other sites. MATERIALS AND METHODS: We reviewed the EHRs of 631 individuals followed at Vanderbilt for hypertension status. We developed features and phenotyping algorithms of increasing complexity. Input categories included International Classification of Diseases, Ninth Revision (ICD9) codes, medications, vital signs, narrative-text search results, and Unified Medical Language System (UMLS) concepts extracted using natural language processing (NLP). We developed a module and tested portability by replicating 10 of the best-performing algorithms at the Marshfield Clinic. RESULTS: Random forests using billing codes, medications, vitals, and concepts had the best performance with a median area under the receiver operator characteristic curve (AUC) of 0.976. Normalized sums of all 4 categories also performed well (0.959 AUC). The best non-NLP algorithm combined normalized ICD9 codes, medications, and blood pressure readings with a median AUC of 0.948. Blood pressure cutoffs or ICD9 code counts alone had AUCs of 0.854 and 0.908, respectively. Marshfield Clinic results were similar. CONCLUSION: This work shows that billing codes or blood pressure readings alone yield good hypertension classification performance. However, even simple combinations of input categories improve performance. The most complex algorithms classified hypertension with excellent recall and precision.
Pedro L. Teixeira, Wei-Qi Wei, Robert M. Cronin, Huan Mo, Jacob P. VanHouten, Robert J. Carroll, Eric LaRose, Lisa Bastarache, S. Trent Rosenbloom, Todd L. Edwards, Dan M. Roden, Thomas A. Lasko, Richard A. Dart, Anne M. Nikolai, Peggy L. Peissig, Joshua C. Denny
J. Am. Medical Informatics Assoc.6
2015 A Genome- and Phenome- Wide Study of Diverticulosis
Yoonjung Y. Joo, Jennifer A. Pacheco, Loren L. Armstrong, William K. Thompson, Robert J. Carroll, Joshua C. Denny, Peggy L. Peissig, James G. Linneman, Jyotishman Pathak, Girish N. Nadkarni, Laura Rasmussen-Torvik, M. Geoffrey Hayes, Abel N. Kho
AMIA5
2014 PheWAS and Genetics Define Subphenotypes in Drug Response
Robert J. Carroll, Jeremy L. Warner, Anne E. Eyler, Charles Moore, Jayanth Doss, Katherine P. Liao, Robert M. Plenge, Joshua C. Denny
AMIA1
2014 Phenome-Wide Association Studies Using NLP-Derived Concepts
Pedro L. Teixeira, Robert J. Carroll, Lisa Bastarache, Peter Speltz, Joshua C. Smith, Joshua C. Denny
AMIA2
2014 R PheWAS: data analysis and plotting tools for phenome-wide association studies in the R environment
abstract
UNLABELLED: Phenome-wide association studies (PheWAS) have been used to replicate known genetic associations and discover new phenotype associations for genetic variants. This PheWAS implementation allows users to translate ICD-9 codes to PheWAS case and control groups, perform analyses using these and/or other phenotypes with covariate adjustments and plot the results. We demonstrate the methods by replicating a PheWAS on rs3135388 (near HLA-DRB, associated with multiple sclerosis) and performing a novel PheWAS using an individual's maximum white blood cell count (WBC) as a continuous measure. Our results for rs3135388 replicate known associations with more significant results than the original study on the same dataset. Our PheWAS of WBC found expected results, including associations with infections, myeloproliferative diseases and associated conditions, such as anemia. These results demonstrate the performance of the improved classification scheme and the flexibility of PheWAS encapsulated in this package. AVAILABILITY AND IMPLEMENTATION: This R package is freely available under the Gnu Public License (GPL-3) from http://phewascatalog.org. It is implemented in native R and is platform independent.
Robert J. Carroll, Lisa Bastarache, Joshua C. Denny
Bioinform.1
2013 Open Source R Implementation of the PheWAS Methodology
Robert J. Carroll, Lisa Bastarache, Joshua C. Denny
AMIA1
2013 Using PheWAS and Natural Language Processing to Discover Clinical Associations for Congenital Chest Deformities
Christine M. McEvoy, Robert J. Carroll, Lisa Bastarache, Wei-Qi Wei, Joshua C. Denny
AMIA2
2012 Using PheWAS to Assess Pleiotropy of Genetic Risk Scores for Rheumatoid Arthritis and Coronary Artery Disease in the eMERGE Network
Robert J. Carroll, Katherine P. Liao, Anne E. Eyler, Lisa Bastarache, Dana C. Crawford, Peggy L. Peissig, Jyotishman Pathak, David Carrell, Abel N. Kho, Rongling Li, Daniel R. Masys, Gail P. Jarvik, Christopher G. Chute, Rex L. Chisholm, Eric B. Larson, Catherine A. McCarty, Iftikhar J. Kullo
AMIA1
2012 Portability of an algorithm to identify rheumatoid arthritis in electronic health records
abstract
OBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining.
Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny
J. Am. Medical Informatics Assoc.1