EDBT 2026 Demo / reviewers in the wild / expert
Jennifer A. Pacheco
dblp:03/11038
· DBLP profile ↗
50ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-8021-5818ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 50 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhenoFit: a framework for determining computable phenotyping algorithm fitness for purpose and reuseabstractBACKGROUND: Computational phenotyping from electronic health records (EHRs) is essential for clinical research, decision support, and quality/population health assessment, but the proliferation of algorithms for the same conditions makes it difficult to identify which algorithm is most appropriate for reuse. OBJECTIVE: To develop a framework for assessing phenotyping algorithm fitness for purpose and reuse. FITNESS FOR PURPOSE: Phenotyping algorithms are fit for purpose when they identify the intended population with performance characteristics appropriate for the intended application. FITNESS FOR REUSE: Phenotyping algorithms are fit for reuse when the algorithm is implementable and generalizable-that is, it identifies the same intended population with similar performance characteristics when applied to a new setting. CONCLUSIONS: The PhenoFit framework provides a structured approach to evaluate and adapt phenotyping algorithms for new contexts increasing efficiency and consistency of identifying patient populations from EHRs. Laura K. Wiley, Luke V. Rasmussen, Rebecca T. Levinson, Jennifer Malinowski, Sheila Manemann, Melissa P. Wilson, Martin Chapman, Jennifer A. Pacheco, Theresa Walunas, Justin Starren, Suzette J. Bielinski, Rachel L. Richesson |
J. Am. Medical Informatics Assoc. | 8 |
| 2023 | Characterizing variability of electronic health record-driven phenotype definitionsabstractOBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic. Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
J. Biomed. Informatics | 6 |
| 2022 | Design and validation of a FHIR-based EHR-driven phenotyping toolboxabstractOBJECTIVES: To develop and validate a standards-based phenotyping tool to author electronic health record (EHR)-based phenotype definitions and demonstrate execution of the definitions against heterogeneous clinical research data platforms. MATERIALS AND METHODS: We developed an open-source, standards-compliant phenotyping tool known as the PhEMA Workbench that enables a phenotype representation using the Fast Healthcare Interoperability Resources (FHIR) and Clinical Quality Language (CQL) standards. We then demonstrated how this tool can be used to conduct EHR-based phenotyping, including phenotype authoring, execution, and validation. We validated the performance of the tool by executing a thrombotic event phenotype definition at 3 sites, Mayo Clinic (MC), Northwestern Medicine (NM), and Weill Cornell Medicine (WCM), and used manual review to determine precision and recall. RESULTS: An initial version of the PhEMA Workbench has been released, which supports phenotype authoring, execution, and publishing to a shared phenotype definition repository. The resulting thrombotic event phenotype definition consisted of 11 CQL statements, and 24 value sets containing a total of 834 codes. Technical validation showed satisfactory performance (both NM and MC had 100% precision and recall and WCM had a precision of 95% and a recall of 84%). CONCLUSIONS: We demonstrate that the PhEMA Workbench can facilitate EHR-driven phenotype definition, execution, and phenotype sharing in heterogeneous clinical research data environments. A phenotype definition that integrates with existing standards-compliant systems, and the use of a formal representation facilitates automation and can decrease potential for human error. Pascal S. Brandt, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Sajjad Abedian, Daniel J. Stone, David Knaack, Jie Xu 0012, Yifan Peng 0002, Natalie C. Benda, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Multi-site Evaluation of Longitudinal Changes in Ejection Fraction in Heart Failure Patients Through Data-driven Phenotyping
Prakash Adekkanattu, Jennifer A. Pacheco, Joseph Kabariti, Daniel J. Stone, Yue Yu 0012, Parag Goyal, Faraz S. Ahmad, Guoqian Jiang, Yuan Luo 0001, Luke V. Rasmussen, Pascal S. Brandt, Jie Xu 0012, Fei Wang 0001, Natalie C. Benda, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 2 |
| 2021 | Supporting EHR-based Cohort Discovery Through User-centered Design: Results of an Early Formative Usability Study
Natalie C. Benda, Pascal S. Brandt, Jessica S. Ancker, Jennifer A. Pacheco, Prakash Adekkanattu, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 4 |
| 2021 | A Deep Learning Framework Using a Pre-trained BERT Model to Predict the Risk of Progression from Mild Cognitive Impairment to Alzheimer's Disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Fei Wang 0001, Richard Isaacson, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 4 |
| 2021 | Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei |
AMIA | 1 |
| 2021 | FHIRTime: Standardizing Temporal Patterns Identified from Clinical Narratives Using HL7 FHIR
Daniel J. Stone, Sijia Liu 0002, Yuan Luo 0001, Andrew Wen, Nansu Zong, Luke V. Rasmussen, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Fei Wang 0001, Cui Tao, Jyotishman Pathak, Guoqian Jiang |
AMIA | 9 |
| 2021 | On Constraints and Considerations for Extending Support for Natural Language Processing-Based FHIR Resource Generation
Andrew Wen, Luke V. Rasmussen, Daniel J. Stone, Sijia Liu 0002, Prakash Adekkanattu, Pascal S. Brandt, Jennifer A. Pacheco, Yuan Luo 0001, Fei Wang 0001, Jyotishman Pathak, Guoqian Jiang |
AMIA | 7 |
| 2020 | Feasibility of Cross-Platform EHR-Driven Phenotyping Using Clinical Quality Language
Pascal S. Brandt, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Evan Sholle, Faraz S. Ahmad, Jie Xu 0012, Jessica S. Ancker, Fei Wang 0001, Yuan Luo 0001, Guoqian Jiang, Jyotishman Pathak, Luke V. Rasmussen |
AMIA | 3 |
| 2020 | Identification of Alzheimer's Disease Subtypes from Electronic Health Records Using a Data-Driven Approach
Jie Xu 0012, Fei Wang 0001, Prakash Adekkanattu, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Yuan Luo 0001, Chengsheng Mao, Jennifer A. Pacheco, Luke V. Rasmussen, Yiye Zhang, Richard Isaacson, Jyotishman Pathak |
AMIA | 10 |
| 2020 | Identifying sub-phenotypes of acute kidney injury using structured and unstructured electronic health record data with memory networks
Jingyuan Chou, Xi Sheryl Zhang, Yuan Luo 0001, Tamara Isakova, Prakash Adekkanattu, Jessica S. Ancker, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Luke V. Rasmussen, Jyotishman Pathak, Fei Wang 0001 |
J. Biomed. Informatics | 10 |
| 2019 | Evaluating the Portability of an NLP System for Processing Echocardiograms: A Retrospective, Multi-site Observational Study
Prakash Adekkanattu, Guoqian Jiang, Yuan Luo 0001, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Richard C. Kiefer, Daniel J. Stone, Pascal S. Brandt, Yizhen Zhong, Fei Wang 0001, Jessica S. Ancker, Thomas R. Campion Jr., Jyotishman Pathak |
AMIA | 7 |
| 2019 | Considerations for Improving the Portability of Electronic Health Record-Based Phenotype Algorithms
Luke V. Rasmussen, Pascal S. Brandt, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Prakash Adekkanattu, Jessica S. Ancker, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
AMIA | 5 |
| 2019 | An ancillary genomics system to support the return of pharmacogenomic resultsabstractExisting approaches to managing genetic and genomic test results from external laboratories typically include filing of text reports within the electronic health record, making them unavailable in many cases for clinical decision support. Even when structured computable results are available, the lack of adopted standards requires considerations for processing the results into actionable knowledge, in addition to storage and management of the data. Here, we describe the design and implementation of an ancillary genomics system used to receive and process heterogeneous results from external laboratories, which returns a descriptive phenotype to the electronic health record in support of pharmacogenetic clinical decision support. Luke V. Rasmussen, Maureen E. Smith, Federico Almaraz, Stephen D. Persell, Laura Rasmussen-Torvik, Jennifer A. Pacheco, Rex L. Chisholm, Carl Christensen, Timothy M. Herr, Firas H. Wehbe, Justin Starren |
J. Am. Medical Informatics Assoc. | 6 |
| 2019 | Developing a FHIR-based EHR phenotyping framework: A case study for identification of patients with obesity and multiple comorbidities from discharge summaries
Na Hong, Andrew Wen, Daniel J. Stone, Shintaro Tsuji, Paul R. Kingsbury, Luke V. Rasmussen, Jennifer A. Pacheco, Prakash Adekkanattu, Fei Wang 0001, Yuan Luo 0001, Jyotishman Pathak, Guoqian Jiang |
J. Biomed. Informatics | 7 |
| 2019 | Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng |
J. Biomed. Informatics | 15 |
| 2018 | Characterizing Design Patterns of EHR-Driven Phenotype Extraction Algorithms
Yizhen Zhong, Luke V. Rasmussen, Jennifer A. Pacheco, Maureen E. Smith, Justin Starren, Wei-Qi Wei, Peter Speltz, Joshua C. Denny, Nephi Walton, George Hripcsak, Christopher G. Chute, Yuan Luo 0001 |
BIBM | 4 |
| 2018 | A case study evaluating the portability of an executable computable phenotype algorithm across multiple institutions and electronic health record environmentsabstractElectronic health record (EHR) algorithms for defining patient cohorts are commonly shared as free-text descriptions that require human intervention both to interpret and implement. We developed the Phenotype Execution and Modeling Architecture (PhEMA, http://projectphema.org) to author and execute standardized computable phenotype algorithms. With PhEMA, we converted an algorithm for benign prostatic hyperplasia, developed for the electronic Medical Records and Genomics network (eMERGE), into a standards-based computable format. Eight sites (7 within eMERGE) received the computable algorithm, and 6 successfully executed it against local data warehouses and/or i2b2 instances. Blinded random chart review of cases selected by the computable algorithm shows PPV ≥90%, and 3 out of 5 sites had >90% overlap of selected cases when comparing the computable algorithm to their original eMERGE implementation. This case study demonstrates potential use of PhEMA computable representations to automate phenotyping across different EHR systems, but also highlights some ongoing challenges. Jennifer A. Pacheco, Luke V. Rasmussen, Richard C. Kiefer, Thomas R. Campion Jr., Peter Speltz, Robert J. Carroll, Sarah C. Stallings, Huan Mo, Monika Ahuja, Guoqian Jiang, Eric LaRose, Peggy L. Peissig, Ning Shang 0004, Barbara Benoit, Vivian S. Gainer, Kenneth Borthwick, Kathryn L. Jackson, Ambrish Sharma, Andy Yizhou Wu, Abel N. Kho, Dan M. Roden, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | A Machine Learning-Based Approach for Identifying Atopic Dermatitis in Adults from Electronic Health Records
Erin N. Gustafson, Al'ona Furmanchuk, Jennifer A. Pacheco, Firas H. Wehbe, Kathryn L. Jackson, Abel N. Kho, William K. Thompson, Jonathan Silverberg |
AMIA | 3 |
| 2017 | Leveraging Value Sets from the Value Set Authority Center (VSAC) in a Standards-Based Clinical Data Repository
Richard C. Kiefer, Luke V. Rasmussen, Jennifer A. Pacheco, Peter Speltz, Joshua C. Denny, William K. Thompson, Jyotishman Pathak, Guoqian Jiang |
AMIA | 3 |
| 2017 | Portable Precision Phenotype Algorithm for Chronic Rhinosinusitis
Jennifer A. Pacheco, Agnes S. Sundaresan, Kenneth Borthwick, Sergio E. Chiarella, David T. Coleman, Andy Yizhou Wu, Abel N. Kho, M. Geoffrey Hayes, Marc S. Williams |
AMIA | 1 |
| 2017 | Sensi-steps: Using Patient-Generated Data to Prevent Post-stroke Falls
Angela Smith, Ada Ng, Eleanor R. Burgess, Jennifer A. Pacheco, Noah D. Weingarten |
AMIA | 4 |
| 2017 | The Phenotype Execution and Modeling Architecture: A Roadmap Towards Next-generation Phenotyping Using EHRs
Peter Speltz, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, William K. Thompson, Guoqian Jiang, Jyotishman Pathak, Joshua C. Denny |
AMIA | 4 |
| 2016 | Personalized Heart Disease Risk Manager: A Tool for Patients and Clinicians to Manage Cardiovascular Risk
Raja Arul Cholan, Jennifer A. Pacheco, Gene Ren, Laura Hickerson |
AMIA | 2 |
| 2016 | An NLP Extension to the Quality Data Model for EHR-Driven Phenotype Algorithm Authoring and Execution
Guoqian Jiang, William K. Thompson, Luke V. Rasmussen, Richard C. Kiefer, Jennifer A. Pacheco, Huan Mo, Peter Speltz, Joshua C. Denny, Jyotishman Pathak |
AMIA | 5 |
| 2016 | Design and Implementation of an Ancillary Genomics System for the Return of Pharmacogenetic Results
Luke V. Rasmussen, Maureen E. Smith, Federico Almaraz, Stephen D. Persell, Laura Rasmussen-Torvik, Jennifer A. Pacheco, Carl Christensen, Timothy M. Herr, Firas H. Wehbe, Justin Starren |
AMIA | 6 |
| 2016 | A multi-institution evaluation of clinical profile anonymizationabstractBACKGROUND AND OBJECTIVE: There is an increasing desire to share de-identified electronic health records (EHRs) for secondary uses, but there are concerns that clinical terms can be exploited to compromise patient identities. Anonymization algorithms mitigate such threats while enabling novel discoveries, but their evaluation has been limited to single institutions. Here, we study how an existing clinical profile anonymization fares at multiple medical centers. METHODS: We apply a state-of-the-artk-anonymization algorithm, withkset to the standard value 5, to the International Classification of Disease, ninth edition codes for patients in a hypothyroidism association study at three medical centers: Marshfield Clinic, Northwestern University, and Vanderbilt University. We assess utility when anonymizing at three population levels: all patients in 1) the EHR system; 2) the biorepository; and 3) a hypothyroidism study. We evaluate utility using 1) changes to the number included in the dataset, 2) number of codes included, and 3) regions generalization and suppression were required. RESULTS: Our findings yield several notable results. First, we show that anonymizing in the context of the entire EHR yields a significantly greater quantity of data by reducing the amount of generalized regions from ∼15% to ∼0.5%. Second, ∼70% of codes that needed generalization only generalized two or three codes in the largest anonymization. CONCLUSIONS: Sharing large volumes of clinical data in support of phenome-wide association studies is possible while safeguarding privacy to the underlying individuals. Raymond Heatherly, Luke V. Rasmussen, Peggy L. Peissig, Jennifer A. Pacheco, Paul A. Harris, Joshua C. Denny, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 4 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 7 |
| 2016 | Developing a data element repository to support EHR-driven phenotype algorithm authoring and execution
Guoqian Jiang, Richard C. Kiefer, Luke V. Rasmussen, Harold R. Solbrig, Huan Mo, Jennifer A. Pacheco, Jie Xu 0011, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
J. Biomed. Informatics | 6 |
| 2015 | Harmonization of Quality Data Model with HL7 FHIR to Support EHR-driven Phenotype Authoring and Execution: A Pilot Study
Guoqian Jiang, Harold R. Solbrig, Richard C. Kiefer, Luke V. Rasmussen, Huan Mo, Jennifer A. Pacheco, Enid N. H. Montague, Jie Xu 0011, Peter Speltz, William K. Thompson, Joshua C. Denny, Christopher G. Chute, Jyotishman Pathak |
AMIA | 6 |
| 2015 | A Genome- and Phenome- Wide Study of Diverticulosis
Yoonjung Y. Joo, Jennifer A. Pacheco, Loren L. Armstrong, William K. Thompson, Robert J. Carroll, Joshua C. Denny, Peggy L. Peissig, James G. Linneman, Jyotishman Pathak, Girish N. Nadkarni, Laura Rasmussen-Torvik, M. Geoffrey Hayes, Abel N. Kho |
AMIA | 2 |
| 2015 | Translating Electronic Clinical Quality Measures to Executable, Portable, and Customizable Workflows in KNIME
Huan Mo, Jennifer A. Pacheco, Richard C. Kiefer, Luke V. Rasmussen, Jyotishman Pathak, Joshua C. Denny, William K. Thompson |
AMIA | 2 |
| 2015 | Usability of a phenotype builder prototype and lessons learned for the design of phenotyping tools
Enid N. H. Montague, Jie Xu 0011, Luke V. Rasmussen, Joshua C. Denny, Guoqian Jiang, Richard C. Kiefer, Jennifer A. Pacheco, Peter Speltz, William K. Thompson, Jyotishman Pathak |
AMIA | 7 |
| 2015 | (Authoring) Rules, (Distributed Query) Tools, and Drools: The challenging new world of high throughput phenotyping
Jennifer A. Pacheco, Abel N. Kho, Jyotishman Pathak, Joshua C. Denny, Shawn N. Murphy |
AMIA | 1 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational researchabstractOBJECTIVE: To review and evaluate available software tools for electronic health record-driven phenotype authoring in order to identify gaps and needs for future development. MATERIALS AND METHODS: Candidate phenotype authoring tools were identified through (1) literature search in four publication databases (PubMed, Embase, Web of Science, and Scopus) and (2) a web search. A collection of tools was compiled and reviewed after the searches. A survey was designed and distributed to the developers of the reviewed tools to discover their functionalities and features. RESULTS: Twenty-four different phenotype authoring tools were identified and reviewed. Developers of 16 of these identified tools completed the evaluation survey (67% response rate). The surveyed tools showed commonalities but also varied in their capabilities in algorithm representation, logic functions, data support and software extensibility, search functions, user interface, and data outputs. DISCUSSION: Positive trends identified in the evaluation included: algorithms can be represented in both computable and human readable formats; and most tools offer a web interface for easy access. However, issues were also identified: many tools were lacking advanced logic functions for authoring complex algorithms; the ability to construct queries that leveraged un-structured data was not widely implemented; and many tools had limited support for plug-ins or external analytic software. CONCLUSIONS: Existing phenotype authoring tools could enable clinical researchers to work with electronic health record data more efficiently, but gaps still exist in terms of the functionalities of such tools. The present work can serve as a reference point for the future development of similar tools. Jie Xu 0011, Luke V. Rasmussen, Pamela L. Shaw, Guoqian Jiang, Richard C. Kiefer, Huan Mo, Jennifer A. Pacheco, Peter Speltz, Qian Zhu 0003, Joshua C. Denny, Jyotishman Pathak, William K. Thompson, Enid N. H. Montague |
J. Am. Medical Informatics Assoc. | 7 |
| 2014 | Automating Extraction and Calculation of Daily Dose and Duration for Medications in EHRs
Jennifer A. Pacheco, William K. Thompson, Kathryn L. Jackson, Abel N. Kho |
AMIA | 1 |
| 2014 | Evaluation of Existing Phenotype Authoring Tools for Clinical Research
Luke V. Rasmussen, Jie Xu 0011, Ruijue Liu, Qian Zhu 0003, Jennifer A. Pacheco, Jyotishman Pathak, William K. Thompson, Joshua C. Denny, Huan Mo, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague |
AMIA | 5 |
| 2014 | Qualitative evaluation of three phenotype information models to find methotrexate liver injury
Qian Zhu 0003, Huan Mo, Luke V. Rasmussen, Andrew R. Post, Jennifer A. Pacheco, Jie Xu 0011, Richard C. Kiefer, Peter Speltz, Enid N. H. Montague, William K. Thompson, Joshua C. Denny, Jyotishman Pathak |
AMIA | 5 |
| 2014 | Replication of SCN5A Associations with Electrocardiographic Traits in African Americans from Clinical and Epidemiologic Studies
Janina M. Jeff, Kristin Brown-Gentry, Robert J. Goodloe, Marylyn D. Ritchie, Joshua C. Denny, Abel N. Kho, Loren L. Armstrong, Bob McClellan Jr., Ping Mayo, Hailing Jin, Niloufar B. Gillani, Nathalie Schnetz-Boutaud, Holli H. Dilks, Melissa A. Basford, Jennifer A. Pacheco, Gail P. Jarvik, Rex L. Chisholm, Dan M. Roden, M. Geoffrey Hayes, Dana C. Crawford |
EvoApplications | 16 |
| 2014 | Design patterns for the development of electronic health record-driven phenotype extraction algorithms
Luke V. Rasmussen, William K. Thompson, Jennifer A. Pacheco, Abel N. Kho, David Carrell, Jyotishman Pathak, Peggy L. Peissig, Gerard Tromp, Joshua C. Denny, Justin Starren |
J. Biomed. Informatics | 3 |
| 2012 | A Geographic Exploration of Colon Polyps
Anna Roberts, Arun Muthalagu, Jennifer A. Pacheco, William K. Thompson, Andrew Gawron, Abel N. Kho |
AMIA | 3 |
| 2012 | An Evaluation of the NQF Quality Data Model for Representing Electronic Health Record Driven Phenotyping Algorithms
William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Peggy L. Peissig, Joshua C. Denny, Abel N. Kho, Aaron W. Miller, Jyotishman Pathak |
AMIA | 3 |
| 2012 | Open Source Workflow Tools for Electronic Health Record Based Phenotyping Algorithms
William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Anna Roberts, Arun Muthalagu, Abel N. Kho |
AMIA | 3 |
| 2012 | Portability of an algorithm to identify rheumatoid arthritis in electronic health recordsabstractOBJECTIVES: Electronic health records (EHR) can allow for the generation of large cohorts of individuals with given diseases for clinical and genomic research. A rate-limiting step is the development of electronic phenotype selection algorithms to find such cohorts. This study evaluated the portability of a published phenotype algorithm to identify rheumatoid arthritis (RA) patients from EHR records at three institutions with different EHR systems. MATERIALS AND METHODS: Physicians reviewed charts from three institutions to identify patients with RA. Each institution compiled attributes from various sources in the EHR, including codified data and clinical narratives, which were searched using one of two natural language processing (NLP) systems. The performance of the published model was compared with locally retrained models. RESULTS: Applying the previously published model from Partners Healthcare to datasets from Northwestern and Vanderbilt Universities, the area under the receiver operating characteristic curve was found to be 92% for Northwestern and 95% for Vanderbilt, compared with 97% at Partners. Retraining the model improved the average sensitivity at a specificity of 97% to 72% from the original 65%. Both the original logistic regression models and locally retrained models were superior to simple billing code count thresholds. DISCUSSION: These results show that a previously published algorithm for RA is portable to two external hospitals using different EHR systems, different NLP systems, and different target NLP vocabularies. Retraining the algorithm primarily increased the sensitivity at each site. CONCLUSION: Electronic phenotype algorithms allow rapid identification of case populations in multiple sites with little retraining. Robert J. Carroll, William K. Thompson, Anne E. Eyler, Arthur M. Mandelin, Tianxi Cai, Raquel M. Zink, Jennifer A. Pacheco, Chad S. Boomershine, Thomas A. Lasko, Hua Xu 0001, Elizabeth W. Karlson, Raúl G. Pérez, Vivian S. Gainer, Shawn N. Murphy, Eric M. Ruderman, Richard M. Pope, Robert M. Plenge, Abel N. Kho, Katherine P. Liao, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 7 |
| 2012 | Use of diverse electronic medical record systems to identify genetic risk for type 2 diabetes within a genome-wide association studyabstractOBJECTIVE: Genome-wide association studies (GWAS) require high specificity and large numbers of subjects to identify genotype-phenotype correlations accurately. The aim of this study was to identify type 2 diabetes (T2D) cases and controls for a GWAS, using data captured through routine clinical care across five institutions using different electronic medical record (EMR) systems. MATERIALS AND METHODS: An algorithm was developed to identify T2D cases and controls based on a combination of diagnoses, medications, and laboratory results. The performance of the algorithm was validated at three of the five participating institutions compared against clinician review. A GWAS was subsequently performed using cases and controls identified by the algorithm, with samples pooled across all five institutions. RESULTS: The algorithm achieved 98% and 100% positive predictive values for the identification of diabetic cases and controls, respectively, as compared against clinician review. By standardizing and applying the algorithm across institutions, 3353 cases and 3352 controls were identified. Subsequent GWAS using data from five institutions replicated the TCF7L2 gene variant (rs7903146) previously associated with T2D. DISCUSSION: By applying stringent criteria to EMR data collected through routine clinical care, cases and controls for a GWAS were identified that subsequently replicated a known genetic variant. The use of standard terminologies to define data elements enabled pooling of subjects and data across five different institutions to achieve the robust numbers required for GWAS. CONCLUSIONS: An algorithm using commonly available data from five different EMR can accurately identify T2D cases and controls for genetic study across multiple institutions. Abel N. Kho, M. Geoffrey Hayes, Laura Rasmussen-Torvik, Jennifer A. Pacheco, William K. Thompson, Loren L. Armstrong, Joshua C. Denny, Peggy L. Peissig, Aaron W. Miller, Wei-Qi Wei, Suzette J. Bielinski, Christopher G. Chute, Cynthia L. Leibson, Gail P. Jarvik, David R. Crosslin, Christopher S. Carlson, Katherine M. Newton, Wendy A. Wolf, Rex L. Chisholm, William L. Lowe |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | Impact of data fragmentation across healthcare centers on the accuracy of a high-throughput clinical phenotyping algorithm for specifying subjects with type 2 diabetes mellitusabstractOBJECTIVE: To evaluate data fragmentation across healthcare centers with regard to the accuracy of a high-throughput clinical phenotyping (HTCP) algorithm developed to differentiate (1) patients with type 2 diabetes mellitus (T2DM) and (2) patients with no diabetes. MATERIALS AND METHODS: This population-based study identified all Olmsted County, Minnesota residents in 2007. We used provider-linked electronic medical record data from the two healthcare centers that provide >95% of all care to County residents (ie, Olmsted Medical Center and Mayo Clinic in Rochester, Minnesota, USA). Subjects were limited to residents with one or more encounter January 1, 2006 through December 31, 2007 at both healthcare centers. DM-relevant data on diagnoses, laboratory results, and medication from both centers were obtained during this period. The algorithm was first executed using data from both centers (ie, the gold standard) and then from Mayo Clinic alone. Positive predictive values and false-negative rates were calculated, and the McNemar test was used to compare categorization when data from the Mayo Clinic alone were used with the gold standard. Age and sex were compared between true-positive and false-negative subjects with T2DM. Statistical significance was accepted as p<0.05. RESULTS: With data from both medical centers, 765 subjects with T2DM (4256 non-DM subjects) were identified. When single-center data were used, 252 T2DM subjects (1573 non-DM subjects) were missed; an additional false-positive 27 T2DM subjects (215 non-DM subjects) were identified. The positive predictive values and false-negative rates were 95.0% (513/540) and 32.9% (252/765), respectively, for T2DM subjects and 92.6% (2683/2898) and 37.0% (1573/4256), respectively, for non-DM subjects. Age and sex distribution differed between true-positive (mean age 62.1; 45% female) and false-negative (mean age 65.0; 56.0% female) T2DM subjects. CONCLUSION: The findings show that application of an HTCP algorithm using data from a single medical center contributes to misclassification. These findings should be considered carefully by researchers when developing and executing HTCP algorithms. Wei-Qi Wei, Cynthia L. Leibson, Jeanine E. Ransom, Abel N. Kho, Pedro J. Caraballo, High Seng Chai, Barbara P. Yawn, Jennifer A. Pacheco, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 8 |
| 2009 | A Highly Specific Algorithm for Identifying Asthma Cases and Controls for Genome-Wide Association Studies
Jennifer A. Pacheco, Pedro C. Avila, Jason A. Thompson, May Law, Jihan A. Quraishi, Alyssa K. Greiman, Eric M. Just, Abel N. Kho |
AMIA | 1 |