VLDB 2026 Research / reviewers in the wild / expert
Melissa A. Basford
dblp:48/8091
· DBLP profile ↗
16ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-0703-2422ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Secondary use of radiological imaging data: Vanderbilt's ImageVU approachabstractOBJECTIVE: To develop ImageVU, a scalable research imaging infrastructure that integrates clinical imaging data with metadata-driven cohort discovery, enabling secure, efficient, and regulatory-compliant access to imaging for secondary and opportunistic research use. This manuscript presents a detailed description of ImageVU's key components and lessons learned to assist other institutions in developing similar research imaging services and infrastructure. METHODS: ImageVU was designed to support the secondary use of radiological imaging data through a dedicated research imaging store. The system comprises four interconnected components: a Research PACS, an Ad Hoc Backfill Host, Cloud Storage System, and a De-Identification System. Imaging metadata are extracted and stored in the Research Derivative (RD), an identified clinical data repository, and the Synthetic Derivative (SD), a de-identified research data repository, with access facilitated through the RD Discover web portal. Researchers interact with the system via structured metadata queries and multiple data delivery options, including web-based viewing, bulk downloads, and dataset preparation for high-performance computing environments. RESULTS: The integration of metadata-driven search capabilities has streamlined cohort discovery and improved imaging data accessibility. As of December 2024, ImageVU has processed 12.9 million MRI and CT series from 1.36 million studies across 453,403 patients. The system has supported 75 project requests, delivering over 50 TB of imaging data to 55 investigators, leading to 66 published research papers. CONCLUSION: ImageVU demonstrates a scalable and efficient approach for integrating clinical imaging into research workflows. By combining institutional data infrastructure with cloud-based storage and metadata-driven cohort identification, the platform enables secure and compliant access to imaging for translational research. David S. Smith, Karthik Ramadass, Laura M. Jones, Jennifer Morse, Daniel Fabbri, Joseph R. Coco, Shunxing Bao, Melissa A. Basford, Peter J. Embí, Reed A. Omary, John C. Gore, Jill M. Pulley, Bennett A. Landman |
J. Biomed. Informatics | 8 |
| 2024 | Empowering the biomedical research community: Innovative SAS deployment on the All of Us Researcher WorkbenchabstractOBJECTIVES: The All of Us Research Program is a precision medicine initiative aimed at establishing a vast, diverse biomedical database accessible through a cloud-based data analysis platform, the Researcher Workbench (RW). Our goal was to empower the research community by co-designing the implementation of SAS in the RW alongside researchers to enable broader use of All of Us data. MATERIALS AND METHODS: Researchers from various fields and with different SAS experience levels participated in co-designing the SAS implementation through user experience interviews. RESULTS: Feedback and lessons learned from user testing informed the final design of the SAS application. DISCUSSION: The co-design approach is critical for reducing technical barriers, broadening All of Us data use, and enhancing the user experience for data analysis on the RW. CONCLUSION: Our co-design approach successfully tailored the implementation of the SAS application to researchers' needs. This approach may inform future software implementations on the RW. Izabelle P. Humes, Cathy Shyr, Moira Dillon, Zhongjie Liu, Jennifer Peterson, Chris De St. Jeor, Jacqueline Malkes, Hiral Master, Brandy Mapes, Romuladus Azuine, Nakia Mack, Bassent Abdelbary, Joyonna Gamble-George, Emily Goldmann, Stephanie Cook, Fatemeh Choupani, Rubin Baskir, Sydney J. McMaster, Chris Lunt, Karriem Watson, Minnkyong Lee, Sophie Schwartz, Ruchi Munshi, David Glazer, Eric Banks, Anthony Philippakis, Melissa A. Basford, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 27 |
| 2024 | Informatics innovation to provide return of value to participant communities in the All of Us Research ProgramabstractOBJECTIVES: The All of Us Research Program harnesses advances in technology, science, and engagement for precision medicine research. We describe informatics innovations which support that goal and return value to the participant cohort and community. MATERIALS AND METHODS: Research data from the All of Us Research Program are available to authorized users on the All of Us Researcher Workbench. We describe the technical infrastructure that enables data access and usage for researchers. Participants are considered partners. To ensure return of value, we outline participant access to information. RESULTS: The All of Us Research Hub allows broad access to data, regardless of background. The innovations described are rooted in the program's core values: participation is open and reflects the diversity of the United States; participants are partners and have access to their information; transparency, security, and privacy are of the highest importance; data are broadly accessible; and the program promotes positive change. We assess research impact and reflect on how All of Us can increase existing return of value to participant communities through future informatics advancements. DISCUSSION: The program will continue to support efforts to ensure equitable access to data and return of value to participants. Looking ahead, we invite the community to join us. CONCLUSION: All of Us research findings can change clinical care, inform guidelines, and set a new bar for data sharing. The ultimate return of value is better care for all. Brandy Mapes, Rachele S. Peterson, Karriem Watson, Melissa A. Basford, Elizabeth Cohn, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | De-black-boxing health AI: demonstrating reproducible machine learning computable phenotypes using the N3C-RECOVER Long COVID model in the All of Us data repositoryabstractMachine learning (ML)-driven computable phenotypes are among the most challenging to share and reproduce. Despite this difficulty, the urgent public health considerations around Long COVID make it especially important to ensure the rigor and reproducibility of Long COVID phenotyping algorithms such that they can be made available to a broad audience of researchers. As part of the NIH Researching COVID to Enhance Recovery (RECOVER) Initiative, researchers with the National COVID Cohort Collaborative (N3C) devised and trained an ML-based phenotype to identify patients highly probable to have Long COVID. Supported by RECOVER, N3C and NIH's All of Us study partnered to reproduce the output of N3C's trained model in the All of Us data enclave, demonstrating model extensibility in multiple environments. This case study in ML-based phenotype reuse illustrates how open-source software best practices and cross-site collaboration can de-black-box phenotyping algorithms, prevent unnecessary rework, and promote open science in informatics. Emily R. Pfaff, Andrew T. Girvin, Miles Crosskey, Srushti Gangireddy, Hiral Master, Wei-Qi Wei, Vern Eric Kerchberger, Mark G. Weiner, Paul A. Harris, Melissa A. Basford, Chris Lunt, Christopher G. Chute, Richard A. Moffitt, Melissa A. Haendel |
J. Am. Medical Informatics Assoc. | 10 |
| 2023 | Managing re-identification risks while providing access to the All of Us research programabstractOBJECTIVE: The All of Us Research Program makes individual-level data available to researchers while protecting the participants' privacy. This article describes the protections embedded in the multistep access process, with a particular focus on how the data was transformed to meet generally accepted re-identification risk levels. METHODS: At the time of the study, the resource consisted of 329 084 participants. Systematic amendments were applied to the data to mitigate re-identification risk (eg, generalization of geographic regions, suppression of public events, and randomization of dates). We computed the re-identification risk for each participant using a state-of-the-art adversarial model specifically assuming that it is known that someone is a participant in the program. We confirmed the expected risk is no greater than 0.09, a threshold that is consistent with guidelines from various US state and federal agencies. We further investigated how risk varied as a function of participant demographics. RESULTS: The results indicated that 95th percentile of the re-identification risk of all the participants is below current thresholds. At the same time, we observed that risk levels were higher for certain race, ethnic, and genders. CONCLUSIONS: While the re-identification risk was sufficiently low, this does not imply that the system is devoid of risk. Rather, All of Us uses a multipronged data protection strategy that includes strong authentication practices, active monitoring of data misuse, and penalization mechanisms for users who violate terms of service. Weiyi Xia, Melissa A. Basford, Robert J. Carroll, Ellen Wright Clayton, Paul A. Harris, Murat Kantarcioglu, Yongtai Liu, Steve Nyemba, Yevgeniy Vorobeychik, Zhiyu Wan, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | The All of Us Research Program Researcher Workbench Phenotype Library: Five Disease Implementations
Izabelle P. Humes, Roxana Loperena-Cortes, Melissa A. Basford, Kelsey R. Mayo, Joseph DiPaolo, David J. Schlueter, Wei-Qi Wei, Robert J. Carroll, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden, Andrea H. Ramirez |
AMIA | 3 |
| 2020 | The All of Us Research Program Researcher Workbench: Cloud based access and analytics to advance precision medicine
Andrea H. Ramirez, Kelsey R. Mayo, Robert J. Carroll, Karthik Muthuraman, Melissa A. Basford, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden |
AMIA | 5 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Replication of SCN5A Associations with Electrocardiographic Traits in African Americans from Clinical and Epidemiologic Studies
Janina M. Jeff, Kristin Brown-Gentry, Robert J. Goodloe, Marylyn D. Ritchie, Joshua C. Denny, Abel N. Kho, Loren L. Armstrong, Bob McClellan Jr., Ping Mayo, Hailing Jin, Niloufar B. Gillani, Nathalie Schnetz-Boutaud, Holli H. Dilks, Melissa A. Basford, Jennifer A. Pacheco, Gail P. Jarvik, Rex L. Chisholm, Dan M. Roden, M. Geoffrey Hayes, Dana C. Crawford |
EvoApplications | 15 |
| 2014 | Brief communication: The Mid-South Clinical Data Research NetworkabstractThe Mid-South Clinical Data Research Network (CDRN) encompasses three large health systems: (1) Vanderbilt Health System (VU) with electronic medical records for over 2 million patients, (2) the Vanderbilt Healthcare Affiliated Network (VHAN) which currently includes over 40 hospitals, hundreds of ambulatory practices, and over 3 million patients in the Mid-South, and (3) Greenway Medical Technologies, with access to 24 million patients nationally. Initial goals of the Mid-South CDRN include: (1) expansion of our VU data network to include the VHAN and Greenway systems, (2) developing data integration/interoperability across the three systems, (3) improving our current tools for extracting clinical data, (4) optimization of tools for collection of patient-reported data, and (5) expansion of clinical decision support. By 18 months, we anticipate our CDRN will robustly support projects in comparative effectiveness research, pragmatic clinical trials, and other key research areas and have the capacity to share data and health information technology tools nationally. S. Trent Rosenbloom, Paul A. Harris, Jill M. Pulley, Melissa A. Basford, Jason Grant, Allison DuBuisson, Russell L. Rothman |
J. Am. Medical Informatics Assoc. | 4 |
| 2014 | Secondary use of clinical data: The Vanderbilt approach
Ioana Danciu, James D. Cowan, Melissa A. Basford, Alexander Saip, Susan Osgood, Jana Shirey-Rice, Jacqueline Kirby, Paul A. Harris |
J. Biomed. Informatics | 3 |
| 2012 | PheKB.org: An Online Collaboration Tool for Phenotype Algorithm Research
Joshua Pruitt, Peter Speltz, Jacqueline Kirby, Melissa A. Basford, Jonathan L. Haines, Joshua C. Denny |
AMIA | 4 |
| 2011 | Mapping clinical phenotype data elements to standardized metadata repositories and controlled terminologies: the eMERGE Network experienceabstractBACKGROUND: Systematic study of clinical phenotypes is important for a better understanding of the genetic basis of human diseases and more effective gene-based disease management. A key aspect in facilitating such studies requires standardized representation of the phenotype data using common data elements (CDEs) and controlled biomedical vocabularies. In this study, the authors analyzed how a limited subset of phenotypic data is amenable to common definition and standardized collection, as well as how their adoption in large-scale epidemiological and genome-wide studies can significantly facilitate cross-study analysis. METHODS: The authors mapped phenotype data dictionaries from five different eMERGE (Electronic Medical Records and Genomics) Network sites studying multiple diseases such as peripheral arterial disease and type 2 diabetes. For mapping, standardized terminological and metadata repository resources, such as the caDSR (Cancer Data Standards Registry and Repository) and SNOMED CT (Systematized Nomenclature of Medicine), were used. The mapping process comprised both lexical (via searching for relevant pre-coordinated concepts and data elements) and semantic (via post-coordination) techniques. Where feasible, new data elements were curated to enhance the coverage during mapping. A web-based application was also developed to uniformly represent and query the mapped data elements from different eMERGE studies. RESULTS: Approximately 60% of the target data elements (95 out of 157) could be mapped using simple lexical analysis techniques on pre-coordinated terms and concepts before any additional curation of terminology and metadata resources was initiated by eMERGE investigators. After curation of 54 new caDSR CDEs and nine new NCI thesaurus concepts and using post-coordination, the authors were able to map the remaining 40% of data elements to caDSR and SNOMED CT. A web-based tool was also implemented to assist in semi-automatic mapping of data elements. CONCLUSION: This study emphasizes the requirement for standardized representation of clinical research data using existing metadata and terminology resources and provides simple techniques and software for data element mapping using experiences from the eMERGE Network. Jyotishman Pathak, Janey Wang, Sudha Kashyap, Melissa A. Basford, Rongling Li, Daniel R. Masys, Christopher G. Chute |
J. Am. Medical Informatics Assoc. | 4 |
| 2011 | Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarinabstractOBJECTIVE: DNA biobanks linked to comprehensive electronic health records systems are potentially powerful resources for pharmacogenetic studies. This study sought to develop natural-language-processing algorithms to extract drug-dose information from clinical text, and to assess the capabilities of such tools to automate the data-extraction process for pharmacogenetic studies. MATERIALS AND METHODS: A manually validated warfarin pharmacogenetic study identified a cohort of 1125 patients with a stable warfarin dose, in which 776 patients were managed by Coumadin Clinic physicians, and the remaining 349 patients were managed by their providers. The authors developed two algorithms to extract weekly warfarin doses from both data sets: a regular expression-based program for semistructured Coumadin Clinic notes; and an advanced weekly dose calculator based on an existing medication information extraction system (MedEx) for narrative providers' notes. The authors then conducted an association analysis between an automatically extracted stable weekly dose of warfarin and four genetic variants of VKORC1 and CYP2C9 genes. The performance of the weekly dose-extraction program was evaluated by comparing it with a gold standard containing manually curated weekly doses. Precision, recall, F-measure, and overall accuracy were reported. Associations between known variants in VKORC1 and CYP2C9 and warfarin stable weekly dose were performed with linear regression adjusted for age, gender, and body mass index. RESULTS: The authors' evaluation showed that the MedEx-based system could determine patients' warfarin weekly doses with 99.7% recall, 90.8% precision, and 93.8% accuracy. Using the automatically extracted weekly doses of warfarin, the authors successfully replicated the previous known associations between warfarin stable dose and genetic variants in VKORC1 and CYP2C9. Hua Xu 0001, Min Jiang 0007, Matthew Oetjens, Erica A. Bowton, Andrea H. Ramirez, Janina M. Jeff, Melissa A. Basford, Jill M. Pulley, James D. Cowan, Marylyn D. Ritchie, Daniel R. Masys, Dan M. Roden, Dana C. Crawford, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 7 |
| 2010 | PheWAS: demonstrating the feasibility of a phenome-wide scan to discover gene-disease associationsabstractMOTIVATION: Emergence of genetic data coupled to longitudinal electronic medical records (EMRs) offers the possibility of phenome-wide association scans (PheWAS) for disease-gene associations. We propose a novel method to scan phenomic data for genetic associations using International Classification of Disease (ICD9) billing codes, which are available in most EMR systems. We have developed a code translation table to automatically define 776 different disease populations and their controls using prevalent ICD9 codes derived from EMR data. As a proof of concept of this algorithm, we genotyped the first 6005 European-Americans accrued into BioVU, Vanderbilt's DNA biobank, at five single nucleotide polymorphisms (SNPs) with previously reported disease associations: atrial fibrillation, Crohn's disease, carotid artery stenosis, coronary artery disease, multiple sclerosis, systemic lupus erythematosus and rheumatoid arthritis. The PheWAS software generated cases and control populations across all ICD9 code groups for each of these five SNPs, and disease-SNP associations were analyzed. The primary outcome of this study was replication of seven previously known SNP-disease associations for these SNPs. RESULTS: Four of seven known SNP-disease associations using the PheWAS algorithm were replicated with P-values between 2.8 x 10(-6) and 0.011. The PheWAS algorithm also identified 19 previously unknown statistical associations between these SNPs and diseases at P < 0.01. This study indicates that PheWAS analysis is a feasible method to investigate SNP-disease associations. Further evaluation is needed to determine the validity of these associations and the appropriate statistical thresholds for clinical significance. AVAILABILITY: The PheWAS software and code translation table are freely available at http://knowledgemap.mc.vanderbilt.edu/research. Joshua C. Denny, Marylyn D. Ritchie, Melissa A. Basford, Jill M. Pulley, Lisa Bastarache, Kristin Brown-Gentry, Deede Wang, Daniel R. Masys, Dan M. Roden, Dana C. Crawford |
Bioinform. | 3 |
| 2010 | An analytical approach to characterize morbidity profile dissimilarity between distinct cohorts using electronic medical records
Jonathan S. Schildcrout, Melissa A. Basford, Jill M. Pulley, Daniel R. Masys, Dan M. Roden, Deede Wang, Christopher G. Chute, Iftikhar J. Kullo, David Carrell, Peggy L. Peissig, Abel N. Kho, Joshua C. Denny |
J. Biomed. Informatics | 2 |