Jill M. Pulley

dblp:69/8091 · DBLP profile ↗
← Back
12ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-0316-5540ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 4 since 2021
YearPublicationVenuePosition
2025 Secondary use of radiological imaging data: Vanderbilt's ImageVU approach
abstract
OBJECTIVE: To develop ImageVU, a scalable research imaging infrastructure that integrates clinical imaging data with metadata-driven cohort discovery, enabling secure, efficient, and regulatory-compliant access to imaging for secondary and opportunistic research use. This manuscript presents a detailed description of ImageVU's key components and lessons learned to assist other institutions in developing similar research imaging services and infrastructure. METHODS: ImageVU was designed to support the secondary use of radiological imaging data through a dedicated research imaging store. The system comprises four interconnected components: a Research PACS, an Ad Hoc Backfill Host, Cloud Storage System, and a De-Identification System. Imaging metadata are extracted and stored in the Research Derivative (RD), an identified clinical data repository, and the Synthetic Derivative (SD), a de-identified research data repository, with access facilitated through the RD Discover web portal. Researchers interact with the system via structured metadata queries and multiple data delivery options, including web-based viewing, bulk downloads, and dataset preparation for high-performance computing environments. RESULTS: The integration of metadata-driven search capabilities has streamlined cohort discovery and improved imaging data accessibility. As of December 2024, ImageVU has processed 12.9 million MRI and CT series from 1.36 million studies across 453,403 patients. The system has supported 75 project requests, delivering over 50 TB of imaging data to 55 investigators, leading to 66 published research papers. CONCLUSION: ImageVU demonstrates a scalable and efficient approach for integrating clinical imaging into research workflows. By combining institutional data infrastructure with cloud-based storage and metadata-driven cohort identification, the platform enables secure and compliant access to imaging for translational research.
David S. Smith, Karthik Ramadass, Laura M. Jones, Jennifer Morse, Daniel Fabbri, Joseph R. Coco, Shunxing Bao, Melissa A. Basford, Peter J. Embí, Reed A. Omary, John C. Gore, Jill M. Pulley, Bennett A. Landman
J. Biomed. Informatics12
2024 PheMIME: an interactive web app and knowledge base for phenome-wide, multi-institutional multimorbidity analysis
abstract
OBJECTIVES: To address the need for interactive visualization tools and databases in characterizing multimorbidity patterns across different populations, we developed the Phenome-wide Multi-Institutional Multimorbidity Explorer (PheMIME). This tool leverages three large-scale EHR systems to facilitate efficient analysis and visualization of disease multimorbidity, aiming to reveal both robust and novel disease associations that are consistent across different systems and to provide insight for enhancing personalized healthcare strategies. MATERIALS AND METHODS: PheMIME integrates summary statistics from phenome-wide analyses of disease multimorbidities, utilizing data from Vanderbilt University Medical Center, Mass General Brigham, and the UK Biobank. It offers interactive and multifaceted visualizations for exploring multimorbidity. Incorporating an enhanced version of associationSubgraphs, PheMIME also enables dynamic analysis and inference of disease clusters, promoting the discovery of complex multimorbidity patterns. A case study on schizophrenia demonstrates its capability for generating interactive visualizations of multimorbidity networks within and across multiple systems. Additionally, PheMIME supports diverse multimorbidity-based discoveries, detailed further in online case studies. RESULTS: The PheMIME is accessible at https://prod.tbilab.org/PheMIME/. A comprehensive tutorial and multiple case studies for demonstration are available at https://prod.tbilab.org/PheMIME_supplementary_materials/. The source code can be downloaded from https://github.com/tbilab/PheMIME. DISCUSSION: PheMIME represents a significant advancement in medical informatics, offering an efficient solution for accessing, analyzing, and interpreting the complex and noisy real-world patient data in electronic health records. CONCLUSION: PheMIME provides an extensive multimorbidity knowledge base that consolidates data from three EHR systems, and it is a novel interactive tool designed to analyze and visualize multimorbidities across multiple EHR datasets. It stands out as the first of its kind to offer extensive multimorbidity knowledge integration with substantial support for efficient online analysis and interactive visualization.
Nick Strayer, Tess Vessels, Karmel Choi, Geoffrey W. Wang, Cosmin Adrian Bejan, Ryan S. Hsi, Alex Bick, Digna R. Velez Edwards, Michael R. Savona, Elizabeth J. Phillips, Jill M. Pulley, Wesley H. Self, Consuelo H. Wilkins, Dan M. Roden, Jordan W. Smoller, Douglas M. Ruderfer, Yaomin Xu
J. Am. Medical Informatics Assoc.13
2023 Interactive network-based clustering and investigation of multimorbidity association matrices with associationSubgraphs
abstract
MOTIVATION: Making sense of networked multivariate association patterns is vitally important to many areas of high-dimensional analysis. Unfortunately, as the data-space dimensions grow, the number of association pairs increases in O(n2); this means that traditional visualizations such as heatmaps quickly become too complicated to parse effectively. RESULTS: Here, we present associationSubgraphs: a new interactive visualization method to quickly and intuitively explore high-dimensional association datasets using network percolation and clustering. The goal is to provide an efficient investigation of association subgraphs, each containing a subset of variables with stronger and more frequent associations among themselves than the remaining variables outside the subset, by showing the entire clustering dynamics and providing subgraphs under all possible cutoff values at once. Particularly, we apply associationSubgraphs to a phenome-wide multimorbidity association matrix generated from an electronic health record and provide an online, interactive demonstration for exploring multimorbidity subgraphs. AVAILABILITY AND IMPLEMENTATION: An R package implementing both the algorithm and visualization components of associationSubgraphs is available at https://github.com/tbilab/associationsubgraphs. Online documentation is available at https://prod.tbilab.org/associationsubgraphs_info/. A demo using a multimorbidity association matrix is available at https://prod.tbilab.org/associationsubgraphs-example/.
Nick Strayer, Lydia Yao, Tess Vessels, Cosmin Adrian Bejan, Ryan S. Hsi, Jana Shirey-Rice, Justin M. Balko, Douglas B. Johnson, Elizabeth J. Phillips, Alex Bick, Todd L. Edwards, Digna R. Velez Edwards, Jill M. Pulley, Quinn Stanton Wells, Michael R. Savona, Nancy J. Cox, Dan M. Roden, Douglas M. Ruderfer, Yaomin Xu
Bioinform.14
2021 PheWAS-ME: a web-app for interactive exploration of multimorbidity patterns in PheWAS
abstract
SUMMARY: Electronic health records (EHRs) linked with a DNA biobank provide unprecedented opportunities for biomedical research in precision medicine. The Phenome-wide association study (PheWAS) is a widely used technique for the evaluation of relationships between genetic variants and a large collection of clinical phenotypes recorded in EHRs. PheWAS analyses are typically presented as static tables and charts of summary statistics obtained from statistical tests of association between a genetic variant and individual phenotypes. Comorbidities are common and typically lead to complex, multivariate gene-disease association signals that are challenging to interpret. Discovering and interrogating multimorbidity patterns and their influence in PheWAS is difficult and time-consuming. We present PheWAS-ME: an interactive dashboard to visualize individual-level genotype and phenotype data side-by-side with PheWAS analysis results, allowing researchers to explore multimorbidity patterns and their associations with a genetic variant of interest. We expect this application to enrich PheWAS analyses by illuminating clinical multimorbidity patterns present in the data. AVAILABILITY AND IMPLEMENTATION: A demo PheWAS-ME application is publicly available at https://prod.tbilab.org/phewas_me/. Sample datasets are provided for exploration with the option to upload custom PheWAS results and corresponding individual-level data. Online versions of the appendices are available at https://prod.tbilab.org/phewas_me_info/. The source code is available as an R package on GitHub (https://github.com/tbilab/multimorbidity_explorer). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Nick Strayer, Jana Shirey-Rice, Shyr Yu, Joshua C. Denny, Jill M. Pulley, Yaomin Xu
Bioinform.5
2018 Mining 100 million notes to find homelessness and adverse childhood experiences: 2 case studies of rare and severe social determinants of health in electronic health records
abstract
Objective: Understanding how to identify the social determinants of health from electronic health records (EHRs) could provide important insights to understand health or disease outcomes. We developed a methodology to capture 2 rare and severe social determinants of health, homelessness and adverse childhood experiences (ACEs), from a large EHR repository. Materials and Methods: We first constructed lexicons to capture homelessness and ACE phenotypic profiles. We employed word2vec and lexical associations to mine homelessness-related words. Next, using relevance feedback, we refined the 2 profiles with iterative searches over 100 million notes from the Vanderbilt EHR. Seven assessors manually reviewed the top-ranked results of 2544 patient visits relevant for homelessness and 1000 patients relevant for ACE. Results: word2vec yielded better performance (area under the precision-recall curve [AUPRC] of 0.94) than lexical associations (AUPRC = 0.83) for extracting homelessness-related words. A comparative study of searches for the 2 phenotypes revealed a higher performance achieved for homelessness (AUPRC = 0.95) than ACE (AUPRC = 0.79). A temporal analysis of the homeless population showed that the majority experienced chronic homelessness. Most ACE patients suffered sexual (70%) and/or physical (50.6%) abuse, with the top-ranked abuser keywords being "father" (21.8%) and "mother" (15.4%). Top prevalent associated conditions for homeless patients were lack of housing (62.8%) and tobacco use disorder (61.5%), while for ACE patients it was mental disorders (36.6%-47.6%). Conclusion: We provide an efficient solution for mining homelessness and ACE information from EHRs, which can facilitate large clinical and genetic studies of these social determinants of health.
Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny
J. Am. Medical Informatics Assoc.8
2017 Large-Scale Text Mining of Social Determinants from Electronic Health Records: Case Studies of Homelessness and Adverse Childhood Experiences
Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny
AMIA8
2015 PheWAS Network Analysis and Visualization
Yaomin Xu, Todd L. Edwards, Lisa Bastarache, Rebecca N. Jerome, Shilin Zho, Eric Torstenson, Wei-Qi Wei, Jana Shirey-Rice, Erica A. Bowton, Shyr Yu, Jill M. Pulley, Joshua C. Denny
AMIA11
2014 Brief communication: The Mid-South Clinical Data Research Network
abstract
The Mid-South Clinical Data Research Network (CDRN) encompasses three large health systems: (1) Vanderbilt Health System (VU) with electronic medical records for over 2 million patients, (2) the Vanderbilt Healthcare Affiliated Network (VHAN) which currently includes over 40 hospitals, hundreds of ambulatory practices, and over 3 million patients in the Mid-South, and (3) Greenway Medical Technologies, with access to 24 million patients nationally. Initial goals of the Mid-South CDRN include: (1) expansion of our VU data network to include the VHAN and Greenway systems, (2) developing data integration/interoperability across the three systems, (3) improving our current tools for extracting clinical data, (4) optimization of tools for collection of patient-reported data, and (5) expansion of clinical decision support. By 18 months, we anticipate our CDRN will robustly support projects in comparative effectiveness research, pragmatic clinical trials, and other key research areas and have the capacity to share data and health information technology tools nationally.
S. Trent Rosenbloom, Paul A. Harris, Jill M. Pulley, Melissa A. Basford, Jason Grant, Allison DuBuisson, Russell L. Rothman
J. Am. Medical Informatics Assoc.3
2011 Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarin
abstract
OBJECTIVE: DNA biobanks linked to comprehensive electronic health records systems are potentially powerful resources for pharmacogenetic studies. This study sought to develop natural-language-processing algorithms to extract drug-dose information from clinical text, and to assess the capabilities of such tools to automate the data-extraction process for pharmacogenetic studies. MATERIALS AND METHODS: A manually validated warfarin pharmacogenetic study identified a cohort of 1125 patients with a stable warfarin dose, in which 776 patients were managed by Coumadin Clinic physicians, and the remaining 349 patients were managed by their providers. The authors developed two algorithms to extract weekly warfarin doses from both data sets: a regular expression-based program for semistructured Coumadin Clinic notes; and an advanced weekly dose calculator based on an existing medication information extraction system (MedEx) for narrative providers' notes. The authors then conducted an association analysis between an automatically extracted stable weekly dose of warfarin and four genetic variants of VKORC1 and CYP2C9 genes. The performance of the weekly dose-extraction program was evaluated by comparing it with a gold standard containing manually curated weekly doses. Precision, recall, F-measure, and overall accuracy were reported. Associations between known variants in VKORC1 and CYP2C9 and warfarin stable weekly dose were performed with linear regression adjusted for age, gender, and body mass index. RESULTS: The authors' evaluation showed that the MedEx-based system could determine patients' warfarin weekly doses with 99.7% recall, 90.8% precision, and 93.8% accuracy. Using the automatically extracted weekly doses of warfarin, the authors successfully replicated the previous known associations between warfarin stable dose and genetic variants in VKORC1 and CYP2C9.
Hua Xu 0001, Min Jiang 0007, Matthew Oetjens, Erica A. Bowton, Andrea H. Ramirez, Janina M. Jeff, Melissa A. Basford, Jill M. Pulley, James D. Cowan, Marylyn D. Ritchie, Daniel R. Masys, Dan M. Roden, Dana C. Crawford, Joshua C. Denny
J. Am. Medical Informatics Assoc.8
2011 StarBRITE: The Vanderbilt University Biomedical Research Integration, Translation and Education portal
Paul A. Harris, Jonathan A. Swafford, Terri L. Edwards, Minhua Zhang, Shraddha S. Nigavekar, Tonya R. Yarbrough, Lynda D. Lane, Tara Helmer, Laurie A. Lebo, Gail Mayo, Daniel R. Masys, Gordon R. Bernard, Jill M. Pulley
J. Biomed. Informatics13
2010 PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene-disease associations
abstract
MOTIVATION: Emergence of genetic data coupled to longitudinal electronic medical records (EMRs) offers the possibility of phenome-wide association scans (PheWAS) for disease-gene associations. We propose a novel method to scan phenomic data for genetic associations using International Classification of Disease (ICD9) billing codes, which are available in most EMR systems. We have developed a code translation table to automatically define 776 different disease populations and their controls using prevalent ICD9 codes derived from EMR data. As a proof of concept of this algorithm, we genotyped the first 6005 European-Americans accrued into BioVU, Vanderbilt's DNA biobank, at five single nucleotide polymorphisms (SNPs) with previously reported disease associations: atrial fibrillation, Crohn's disease, carotid artery stenosis, coronary artery disease, multiple sclerosis, systemic lupus erythematosus and rheumatoid arthritis. The PheWAS software generated cases and control populations across all ICD9 code groups for each of these five SNPs, and disease-SNP associations were analyzed. The primary outcome of this study was replication of seven previously known SNP-disease associations for these SNPs. RESULTS: Four of seven known SNP-disease associations using the PheWAS algorithm were replicated with P-values between 2.8 x 10(-6) and 0.011. The PheWAS algorithm also identified 19 previously unknown statistical associations between these SNPs and diseases at P < 0.01. This study indicates that PheWAS analysis is a feasible method to investigate SNP-disease associations. Further evaluation is needed to determine the validity of these associations and the appropriate statistical thresholds for clinical significance. AVAILABILITY: The PheWAS software and code translation table are freely available at http://knowledgemap.mc.vanderbilt.edu/research.
Joshua C. Denny, Marylyn D. Ritchie, Melissa A. Basford, Jill M. Pulley, Lisa Bastarache, Kristin Brown-Gentry, Deede Wang, Daniel R. Masys, Dan M. Roden, Dana C. Crawford
Bioinform.4
2010 An analytical approach to characterize morbidity profile dissimilarity between distinct cohorts using electronic medical records
Jonathan S. Schildcrout, Melissa A. Basford, Jill M. Pulley, Daniel R. Masys, Dan M. Roden, Deede Wang, Christopher G. Chute, Iftikhar J. Kullo, David Carrell, Peggy L. Peissig, Abel N. Kho, Joshua C. Denny
J. Biomed. Informatics3