EDBT 2026 Demo / reviewers in the wild / expert
Paul A. Harris
dblp:11/7390
· DBLP profile ↗
42ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-1744-2011ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 42 · 6 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multisite evaluation of automated electronic case report form data entry from electronic health recordsabstractOBJECTIVE: Multicenter clinical trials often abstract data from the electronic health record (EHR) onto a case report form (CRF) via an electronic data capture (EDC) system. The abstraction process is manual, time-consuming, and error prone. We evaluated scaling automated CRF completion from one institution to other sites in a multicenter trial. METHODS: We exported a REDCap project with embedded EHR mapping for a completed platform trial from one institution and delivered it to two other study sites. The receiving sites determined whether additional data elements could be mapped for their institution. We measured the proportion of data entry that could be automated, the extent of agreement between the human- and automation-entered data, and the staff effort required to set up automated CRF completion. RESULTS: It took approximately 26 and 15 h to set up automation and to map data from the EHR systems at the two receiving institutions, respectively. For 20 total participants at the two receiving institutions, out of 4404 fields with human-entered data, using CDIS could have prevented 764 data entry errors that persisted after monitoring and would have saved 17 total hours or 51 min per participant of manual data entry time. CONCLUSION: With these initial estimates of the configuration and data re-mapping time required to scale automated CRF completion and impact on data quality, investigators planning multicenter trials are better positioned to determine when benefits of automation outweigh the expense of manual data abstraction for clinical trials. Alex C. Cheng, M. Katie Banasiewicz, Kevin W. Gibbs, Genesis Briceno, Dena Iadanza, Akram Khan, Leigha Landreth, Bas de Veer, Elizabeth L. Moyer, Kevin P. Seitz, Jakea D. Johnson, Francesco Delacqua, Adam A. Lewis, Sean P. Collins, Wesley H. Self, Matthew S. Shotwell, Christopher J. Lindsell, Jonathan D. Casey, Paul A. Harris |
J. Biomed. Informatics | 19 |
| 2025 | Supporting rapid innovation in research data capture and management: the REDCap external module frameworkabstractOBJECTIVES: Establishing a robust and secure framework allowing creation and sharing of custom features within the REDCap electronic data capture platform. MATERIALS AND METHODS: In partnership with REDCap Consortium members, we developed a framework for creating external modules enabling project-specific REDCap custom functionality (EM Framework). The EM Framework includes guidance and standard processes for developers to ensure basic functionality, compatibility, and security across REDCap instances. The EM Framework also includes an optional dissemination mechanism, the REDCap Repository of External Modules (Repo), for developers to easily share their work with other institutions in the REDCap Consortium. RESULTS: From the EM Framework's launch in 2017 through 2024, 356 external modules have been published to the Repo by software developers at 59 institutions. These modules have been used on 29 485 projects at 2107 institutions in 67 countries. Over time, features from 22 of these external modules have been integrated into the core REDCap code serving 7700+ REDCap Consortium members in 160 countries. DISCUSSION: The EM Framework permits developers to create, test, and deploy custom features to their local REDCap platform. It further enables a process to distribute these features to other REDCap administrators across the Consortium. CONCLUSION: The EM Framework has enhanced innovation in electronic data capture and dissemination of those innovations to a global research community. Alex C. Cheng, Stephany N. Duda, Kyle McGuffin, Mark McEver, Robert Taylor 0001, Günther A. Rezniczek, Eduardo Morales, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | Reply to Layne et al.'s Letter to the EditorabstractWe appreciate Layne et al.’s comments regarding our recent study on leveraging AI to generate lay summaries of scientific abstracts.1 We agree with the authors that readability is a key aspect of writing effective lay summaries. The authors noted that we “prompted ChatGPT-4 to craft a lay summary ‘in lay language at a 6th grade reading level’ with no other focus in the prompt on readability or other suggestions by the American Medical Association (AMA) recommendations for lay summaries.”2 In response to this, we would like to respectfully point out that our prompt also emphasizes succinctness (“under 100 words”) and clear focus on the key components of a scientific abstract (“highlight the study purpose, methods, key findings, and practical importance of these findings”). These elements are critical to readability and align with the AMA’s checklist for creating written materials for a lay audience.2 We acknowledge that readability formulas (eg, Flesch–Kincaid readability score, SMOG Index) can serve as useful tools for assessing the difficulty of the vocabulary and sentences in lay summaries. However, it is well known that these formulas overlook important factors that influence ease of reading, including content and the reader’s prior knowledge; as a result, their assessments can be inconsistent and often inaccurate.3,4 In the 2020 US Department of Health and Human Services’ guidance document on using readability formulas, it is cautioned that “relying on a grade level score can mislead you into thinking that your materials are clear and effective when they are not.”5 These limitations underscore the need for more comprehensive methods to evaluate the effectiveness of AI-generated content for lay audiences. As AI’s vision and language capabilities continue to evolve, there is potential to leverage these advancements to generate multimodal lay summaries, including AI-generated illustrations to supplement written information and translations into multiple languages. Consequently, evaluation methods would need to evolve and adapt to rigorously assess multimodal content designed for diverse audiences. Key considerations include assessing the accuracy, clarity, potential harm, and cultural relevance to the target audience to better understand the real-world impact of AI-generated materials on public comprehension and engagement with scientific results. Cathy Shyr, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 2 |
| 2025 | A REDCap advanced randomization module to meet the needs of modern trials
Luke Stevens, Nan Kennedy, Robert J. Taylor, Adam A. Lewis, Frank E. Harrell, Matthew S. Shotwell, Emily S. Serdoz, Gordon R. Bernard, Wesley H. Self, Christopher J. Lindsell, Paul A. Harris, Jonathan D. Casey |
J. Biomed. Informatics | 11 |
| 2024 | Empowering the biomedical research community: Innovative SAS deployment on the All of Us Researcher WorkbenchabstractOBJECTIVES: The All of Us Research Program is a precision medicine initiative aimed at establishing a vast, diverse biomedical database accessible through a cloud-based data analysis platform, the Researcher Workbench (RW). Our goal was to empower the research community by co-designing the implementation of SAS in the RW alongside researchers to enable broader use of All of Us data. MATERIALS AND METHODS: Researchers from various fields and with different SAS experience levels participated in co-designing the SAS implementation through user experience interviews. RESULTS: Feedback and lessons learned from user testing informed the final design of the SAS application. DISCUSSION: The co-design approach is critical for reducing technical barriers, broadening All of Us data use, and enhancing the user experience for data analysis on the RW. CONCLUSION: Our co-design approach successfully tailored the implementation of the SAS application to researchers' needs. This approach may inform future software implementations on the RW. Izabelle P. Humes, Cathy Shyr, Moira Dillon, Zhongjie Liu, Jennifer Peterson, Chris De St. Jeor, Jacqueline Malkes, Hiral Master, Brandy Mapes, Romuladus Azuine, Nakia Mack, Bassent Abdelbary, Joyonna Gamble-George, Emily Goldmann, Stephanie Cook, Fatemeh Choupani, Rubin Baskir, Sydney J. McMaster, Chris Lunt, Karriem Watson, Minnkyong Lee, Sophie Schwartz, Ruchi Munshi, David Glazer, Eric Banks, Anthony Philippakis, Melissa A. Basford, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 29 |
| 2024 | Informatics innovation to provide return of value to participant communities in the All of Us Research ProgramabstractOBJECTIVES: The All of Us Research Program harnesses advances in technology, science, and engagement for precision medicine research. We describe informatics innovations which support that goal and return value to the participant cohort and community. MATERIALS AND METHODS: Research data from the All of Us Research Program are available to authorized users on the All of Us Researcher Workbench. We describe the technical infrastructure that enables data access and usage for researchers. Participants are considered partners. To ensure return of value, we outline participant access to information. RESULTS: The All of Us Research Hub allows broad access to data, regardless of background. The innovations described are rooted in the program's core values: participation is open and reflects the diversity of the United States; participants are partners and have access to their information; transparency, security, and privacy are of the highest importance; data are broadly accessible; and the program promotes positive change. We assess research impact and reflect on how All of Us can increase existing return of value to participant communities through future informatics advancements. DISCUSSION: The program will continue to support efforts to ensure equitable access to data and return of value to participants. Looking ahead, we invite the community to join us. CONCLUSION: All of Us research findings can change clinical care, inform guidelines, and set a new bar for data sharing. The ultimate return of value is better care for all. Brandy Mapes, Rachele S. Peterson, Karriem Watson, Melissa A. Basford, Elizabeth Cohn, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Leveraging artificial intelligence to summarize abstracts in lay language for increasing research accessibility and transparencyabstractOBJECTIVE: Returning aggregate study results is an important ethical responsibility to promote trust and inform decision making, but the practice of providing results to a lay audience is not widely adopted. Barriers include significant cost and time required to develop lay summaries and scarce infrastructure necessary for returning them to the public. Our study aims to generate, evaluate, and implement ChatGPT 4 lay summaries of scientific abstracts on a national clinical study recruitment platform, ResearchMatch, to facilitate timely and cost-effective return of study results at scale. MATERIALS AND METHODS: We engineered prompts to summarize abstracts at a literacy level accessible to the public, prioritizing succinctness, clarity, and practical relevance. Researchers and volunteers assessed ChatGPT-generated lay summaries across five dimensions: accuracy, relevance, accessibility, transparency, and harmfulness. We used precision analysis and adaptive random sampling to determine the optimal number of summaries for evaluation, ensuring high statistical precision. RESULTS: ChatGPT achieved 95.9% (95% CI, 92.1-97.9) accuracy and 96.2% (92.4-98.1) relevance across 192 summary sentences from 33 abstracts based on researcher review. 85.3% (69.9-93.6) of 34 volunteers perceived ChatGPT-generated summaries as more accessible and 73.5% (56.9-85.4) more transparent than the original abstract. None of the summaries were deemed harmful. We expanded ResearchMatch's technical infrastructure to automatically generate and display lay summaries for over 750 published studies that resulted from the platform's recruitment mechanism. DISCUSSION AND CONCLUSION: Implementing AI-generated lay summaries on ResearchMatch demonstrates the potential of a scalable framework generalizable to broader platforms for enhancing research accessibility and transparency. Cathy Shyr, Randall W. Grout, Nan Kennedy, Yasemin Akdas, Maeve Tischbein, Joshua Milford, Jason Tan, Kaysi Quarles, Terri L. Edwards, Laurie L. Novak, Jules White, Consuelo H. Wilkins, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 13 |
| 2024 | Illuminating the landscape of high-level clinical trial opportunities in the All of Us Research ProgramabstractOBJECTIVE: With its size and diversity, the All of Us Research Program has the potential to power and improve representation in clinical trials through ancillary studies like Nutrition for Precision Health. We sought to characterize high-level trial opportunities for the diverse participants and sponsors of future trial investment. MATERIALS AND METHODS: We matched All of Us participants with available trials on ClinicalTrials.gov based on medical conditions, age, sex, and geographic location. Based on the number of matched trials, we (1) developed the Trial Opportunities Compass (TOC) to help sponsors assess trial investment portfolios, (2) characterized the landscape of trial opportunities in a phenome-wide association study (PheWAS), and (3) assessed the relationship between trial opportunities and social determinants of health (SDoH) to identify potential barriers to trial participation. RESULTS: Our study included 181 529 All of Us participants and 18 634 trials. The TOC identified opportunities for portfolio investment and gaps in currently available trials across federal, industrial, and academic sponsors. PheWAS results revealed an emphasis on mental disorder-related trials, with anxiety disorder having the highest adjusted increase in the number of matched trials (59% [95% CI, 57-62]; P < 1e-300). Participants from certain communities underrepresented in biomedical research, including self-reported racial and ethnic minorities, had more matched trials after adjusting for other factors. Living in a nonmetropolitan area was associated with up to 13.1 times fewer matched trials. DISCUSSION AND CONCLUSION: All of Us data are a valuable resource for identifying trial opportunities to inform trial portfolio planning. Characterizing these opportunities with consideration for SDoH can provide guidance on prioritizing the most pressing barriers to trial participation. Cathy Shyr, Lina M. Sulieman, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 3 |
| 2024 | Identifying erroneous height and weight values from adult electronic health records in the All of Us research programabstractINTRODUCTION: Electronic Health Records (EHR) are a useful data source for research, but their usability is hindered by measurement errors. This study investigated an automatic error detection algorithm for adult height and weight measurements in EHR for the All of Us Research Program (All of Us). METHODS: We developed reference charts for adult heights and weights that were stratified on participant sex. Our analysis included 4,076,534 height and 5,207,328 wt measurements from ∼ 150,000 participants. Errors were identified using modified standard deviation scores, differences from their expected values, and significant changes between consecutive measurements. We evaluated our method with chart-reviewed heights (8,092) and weights (9,039) from 250 randomly selected participants and compared it with the current cleaning algorithm in All of Us. RESULTS: The proposed algorithm classified 1.4 % of height and 1.5 % of weight errors in the full cohort. Sensitivity was 90.4 % (95 % CI: 79.0-96.8 %) for heights and 65.9 % (95 % CI: 56.9-74.1 %) for weights. Precision was 73.4 % (95 % CI: 60.9-83.7 %) for heights and 62.9 (95 % CI: 54.0-71.1 %) for weights. In comparison, the current cleaning algorithm has inferior performance in sensitivity (55.8 %) and precision (16.5 %) for height errors while having higher precision (94.0 %) and lower sensitivity (61.9 %) for weight errors. DISCUSSION: Our proposed algorithm outperformed in detecting height errors compared to weights. It can serve as a valuable addition to the current All of Us cleaning algorithm for identifying erroneous height values. Andrew Guide, Lina M. Sulieman, Shawn Garbett, Robert M. Cronin, Matthew E. Spotnitz, Karthik Natarajan, Robert J. Carroll, Paul A. Harris, Qingxia Chen |
J. Biomed. Informatics | 8 |
| 2023 | De-black-boxing health AI: demonstrating reproducible machine learning computable phenotypes using the N3C-RECOVER Long COVID model in the All of Us data repositoryabstractMachine learning (ML)-driven computable phenotypes are among the most challenging to share and reproduce. Despite this difficulty, the urgent public health considerations around Long COVID make it especially important to ensure the rigor and reproducibility of Long COVID phenotyping algorithms such that they can be made available to a broad audience of researchers. As part of the NIH Researching COVID to Enhance Recovery (RECOVER) Initiative, researchers with the National COVID Cohort Collaborative (N3C) devised and trained an ML-based phenotype to identify patients highly probable to have Long COVID. Supported by RECOVER, N3C and NIH's All of Us study partnered to reproduce the output of N3C's trained model in the All of Us data enclave, demonstrating model extensibility in multiple environments. This case study in ML-based phenotype reuse illustrates how open-source software best practices and cross-site collaboration can de-black-box phenotyping algorithms, prevent unnecessary rework, and promote open science in informatics. Emily R. Pfaff, Andrew T. Girvin, Miles Crosskey, Srushti Gangireddy, Hiral Master, Wei-Qi Wei, Vern Eric Kerchberger, Mark G. Weiner, Paul A. Harris, Melissa A. Basford, Chris Lunt, Christopher G. Chute, Richard A. Moffitt, Melissa A. Haendel |
J. Am. Medical Informatics Assoc. | 9 |
| 2023 | Managing re-identification risks while providing access to the All of Us research programabstractOBJECTIVE: The All of Us Research Program makes individual-level data available to researchers while protecting the participants' privacy. This article describes the protections embedded in the multistep access process, with a particular focus on how the data was transformed to meet generally accepted re-identification risk levels. METHODS: At the time of the study, the resource consisted of 329 084 participants. Systematic amendments were applied to the data to mitigate re-identification risk (eg, generalization of geographic regions, suppression of public events, and randomization of dates). We computed the re-identification risk for each participant using a state-of-the-art adversarial model specifically assuming that it is known that someone is a participant in the program. We confirmed the expected risk is no greater than 0.09, a threshold that is consistent with guidelines from various US state and federal agencies. We further investigated how risk varied as a function of participant demographics. RESULTS: The results indicated that 95th percentile of the re-identification risk of all the participants is below current thresholds. At the same time, we observed that risk levels were higher for certain race, ethnic, and genders. CONCLUSIONS: While the re-identification risk was sufficiently low, this does not imply that the system is devoid of risk. Rather, All of Us uses a multipronged data protection strategy that includes strong authentication practices, active monitoring of data misuse, and penalization mechanisms for users who violate terms of service. Weiyi Xia, Melissa A. Basford, Robert J. Carroll, Ellen Wright Clayton, Paul A. Harris, Murat Kantarcioglu, Yongtai Liu, Steve Nyemba, Yevgeniy Vorobeychik, Zhiyu Wan, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | Self-paced Training Modality to Promote the Use of All of Us Researcher Workbench in Educational and Research Settings
Hiral Master, Lina M. Sulieman, Paul A. Harris, Karthik Natarajan, Robert J. Carroll, Kayla Marginean, Kelsey R. Mayo, Aymone Kouame |
AMIA | 3 |
| 2022 | Assessing Data Quality and Diversity in the All of Us
Lina M. Sulieman, Jennifer Zhang, Kayla Marginean, Paul A. Harris, Robert J. Carroll |
AMIA | 4 |
| 2022 | HL7 FHIR-based tools and initiatives to support clinical research: a scoping reviewabstractOBJECTIVES: The HL7® fast healthcare interoperability resources (FHIR®) specification has emerged as the leading interoperability standard for the exchange of healthcare data. We conducted a scoping review to identify trends and gaps in the use of FHIR for clinical research. MATERIALS AND METHODS: We reviewed published literature, federally funded project databases, application websites, and other sources to discover FHIR-based papers, projects, and tools (collectively, "FHIR projects") available to support clinical research activities. RESULTS: Our search identified 203 different FHIR projects applicable to clinical research. Most were associated with preparations to conduct research, such as data mapping to and from FHIR formats (n = 66, 32.5%) and managing ontologies with FHIR (n = 30, 14.8%), or post-study data activities, such as sharing data using repositories or registries (n = 24, 11.8%), general research data sharing (n = 23, 11.3%), and management of genomic data (n = 21, 10.3%). With the exception of phenotyping (n = 19, 9.4%), fewer FHIR-based projects focused on needs within the clinical research process itself. DISCUSSION: Funding and usage of FHIR-enabled solutions for research are expanding, but most projects appear focused on establishing data pipelines and linking clinical systems such as electronic health records, patient-facing data systems, and registries, possibly due to the relative newness of FHIR and the incentives for FHIR integration in health information systems. Fewer FHIR projects were associated with research-only activities. CONCLUSION: The FHIR standard is becoming an essential component of the clinical research enterprise. To develop FHIR's full potential for clinical research, funding and operational stakeholders should address gaps in FHIR-based research tools and methods. Stephany Duda, Nan Kennedy, Douglas Conway, Alex C. Cheng, Viet Nguyen, Teresa Zayas-Cabán, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 7 |
| 2022 | EHR-based cohort assessment for multicenter RCTs: a fast and flexible model for identifying potential study sitesabstractOBJECTIVE: The Recruitment Innovation Center (RIC), partnering with the Trial Innovation Network and institutions in the National Institutes of Health-sponsored Clinical and Translational Science Awards (CTSA) Program, aimed to develop a service line to retrieve study population estimates from electronic health record (EHR) systems for use in selecting enrollment sites for multicenter clinical trials. Our goal was to create and field-test a low burden, low tech, and high-yield method. MATERIALS AND METHODS: In building this service line, the RIC strove to complement, rather than replace, CTSA hubs' existing cohort assessment tools. For each new EHR cohort request, we work with the investigator to develop a computable phenotype algorithm that targets the desired population. CTSA hubs run the phenotype query and return results using a standardized survey. We provide a comprehensive report to the investigator to assist in study site selection. RESULTS: From 2017 to 2020, the RIC developed and socialized 36 phenotype-dependent cohort requests on behalf of investigators. The average response rate to these requests was 73%. DISCUSSION: Achieving enrollment goals in a multicenter clinical trial requires that researchers identify study sites that will provide sufficient enrollment. The fast and flexible method the RIC has developed, with CTSA feedback, allows hubs to query their EHR using a generalizable, vetted phenotype algorithm to produce reliable counts of potentially eligible study participants. CONCLUSION: The RIC's EHR cohort assessment process for evaluating sites for multicenter trials has been shown to be efficient and helpful. The model may be replicated for use by other programs. Sarah J. Nelson, Bethany Drury, Daniel Hood, Jeremy Harper, Tiffany Bernard, Chunhua Weng, Nan Kennedy, Bernard LaSalle, Ramkiran Gouripeddi, Consuelo H. Wilkins, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 11 |
| 2022 | Comparing medical history data derived from electronic health records and survey answers in the All of Us Research ProgramabstractOBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases. Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 8 |
| 2021 | Evaluating HL7 FHIR Resources for Sharing Research Consent Data
M. Katie Banasiewicz, Mark McEver, Douglas Conway, Alex C. Cheng, Colleen Lawrence, Leah Dunkel, Paul A. Harris, Stephany Duda |
AMIA | 7 |
| 2021 | Data Coordination for Multi-Site Clinical Trials Using the REDCap Application Programming Interface
Alex C. Cheng, Mark McEver, Francesco Delacqua, Adam A. Lewis, Patrick Newman, Paul A. Harris |
AMIA | 6 |
| 2021 | Promoting Use of Common Data Elements in Research Studies
Paul A. Harris, Robert J. Taylor, Vaishali Jagtap, Douglas Conway, Stephany Duda, Alex C. Cheng |
AMIA | 1 |
| 2021 | Measuring the correctness of All of Us physical measurement
Lina M. Sulieman, Karthik Natarajan, Qingxia Chen, Robert J. Carroll, Kayla Marginean, Paul A. Harris, Andrea H. Ramirez |
AMIA | 6 |
| 2021 | Advancing the Use of FHIR in Research: An Update on NIH's Efforts
Teresa Zayas-Cabán, Belinda Seto, Paul A. Harris, Allison P. Heath, Viet Nguyen |
AMIA | 3 |
| 2021 | REDCap on FHIR: Clinical Data Interoperability Services
Alex C. Cheng, Stephany N. Duda, Robert Taylor 0001, Francesco Delacqua, Adam A. Lewis, Teresa Bosler, Kevin B. Johnson, Paul A. Harris |
J. Biomed. Informatics | 8 |
| 2021 | Creating and implementing a COVID-19 recruitment Data MartabstractThe COVID-19 pandemic has resulted in an unprecedented strain on every aspect of the healthcare system, and clinical research is no exception. Researchers are working against the clock to ramp up research studies addressing every angle of COVID-19 - gaining a better understanding of person-to-person transmission, improving methods for diagnosis, and developing therapies to treat infection and vaccines to prevent it. The impact of the virus on research efforts is not limited to investigators and their teams. Potential participants also face unparalleled opportunities and requests to participate in research, which can result in a significant amount of participant fatigue. The Vanderbilt Institute for Clinical and Translational Research recognized early in the pandemic that a solution to assist researchers in the rapid identification of potential participants was critical, and thus developed the COVID-19 Recruitment Data Mart. This solution does not rest solely on technology; the addition of experienced project managers to support researchers and facilitate collaboration was essential. Since the platform and study support tools were launched on July 20, 2020, four studies have been onboarded and a total of 1693 potential participant matches have been shared. Each of these patients had agreed in advance to direct contact for COVID-19 research and had been matched to study-specific inclusion/exclusion criteria. Our innovative Data Mart system is scalable and looks promising as a generalizable solution for simultaneously recommending individuals from a pool of patients against a pool of time-sensitive trial opportunities. Tara Helmer, Adam A. Lewis, Mark McEver, Francesco Delacqua, Cindy L. Pastern, Nan Kennedy, Terri L. Edwards, Beverly O. Woodward, Paul A. Harris |
J. Biomed. Informatics | 9 |
| 2020 | Implementation of the HL7 FHIR Questionnaire Resource for REDCap
Stephany Duda, Mark McEver, Douglas Conway, Teresa Zayas-Cabán, Paul A. Harris |
AMIA | 5 |
| 2020 | Fitbit "Bring Your Own Device" data in the All of Us Research Program
Michelle Holko, Francis Ratsimbazafy, Kayla Marginean, Karthik Natarajan, Sylvia Cho, Josh Schilling, Aymone Kouame, Dan Webster, Shaquille Peters, Mark Begale, Kelly Gebo, Andrea H. Ramirez, Paul A. Harris |
AMIA | 13 |
| 2020 | The All of Us Research Program Researcher Workbench Phenotype Library: Five Disease Implementations
Izabelle P. Humes, Roxana Loperena-Cortes, Melissa A. Basford, Kelsey R. Mayo, Joseph DiPaolo, David J. Schlueter, Wei-Qi Wei, Robert J. Carroll, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden, Andrea H. Ramirez |
AMIA | 11 |
| 2020 | The All of Us Research Program Researcher Workbench: Cloud based access and analytics to advance precision medicine
Andrea H. Ramirez, Kelsey R. Mayo, Robert J. Carroll, Karthik Muthuraman, Melissa A. Basford, David Glazer, Paul A. Harris, Anthony A. Philippakis, Dan M. Roden |
AMIA | 7 |
| 2019 | The REDCap consortium: Building an international community of software platform partners
Paul A. Harris, Robert Taylor 0001, Brenda L. Minor, Veida Elliott, Michelle Fernandez, Lindsay O'Neal, Laura McLeod, Giovanni Delacqua, Francesco Delacqua, Jacqueline Kirby, Stephany N. Duda |
J. Biomed. Informatics | 1 |
| 2018 | Patient and healthcare provider views on a patient-reported outcomes portalabstractBackground: Over the past decade, public interest in managing health-related information for personal understanding and self-improvement has rapidly expanded. This study explored aspects of how patient-provided health information could be obtained through an electronic portal and presented to inform and engage patients while also providing information for healthcare providers. Methods: We invited participants using ResearchMatch from 2 cohorts: (1) self-reported healthy volunteers (no medical conditions) and (2) individuals with a self-reported diagnosis of anxiety and/or depression. Participants used a secure web application (dashboard) to complete the PROMIS® domain survey(s) and then complete a feedback survey. A community engagement studio with 5 healthcare providers assessed perspectives on the feasibility and features of a portal to collect and display patient provided health information. We used bivariate analyses and regression analyses to determine differences between cohorts. Results: A total of 480 participants completed the study (239 healthy, 241 anxiety and/or depression). While participants from the tw2o cohorts had significantly different PROMIS scores (p < .05), both cohorts welcomed the concept of a patient-centric dashboard, saw value in sharing results with their healthcare provider, and wanted to view results over time. However, factors needing consideration before widespread use included personalization for the patient and their health issues, integration with existing information (eg electronic health records), and integration into clinician workflow. Conclusions: Our findings demonstrated a strong desire among healthy people, patients with chronic diseases, and healthcare providers for a self-assessment portal that can collect patient-reported outcome metrics and deliver personalized feedback. Robert M. Cronin, Douglas Conway, David M. Condon, Rebecca N. Jerome, Daniel W. Byrne, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 6 |
| 2017 | HealthPro: An integrated web application for essential health data and biological specimen collection in the Precision Medicine Initiative
Kelsey R. Mayo, Robert J. Carroll, Jason Tan, Rebecca Johnston, Celecia M. Scott, Joshua C. Denny, Paul A. Harris |
AMIA | 7 |
| 2017 | The All of Us Research Program Researcher Portal: Innovative access to Unprecendented Data
Andrea H. Ramirez, Anthony A. Philippakis, Gonçalo R. Abecasis, Paul A. Harris, Joshua C. Denny |
AMIA | 4 |
| 2016 | Assessment Center API: A Software Component Model for the Integration of Patient Reported Outcomes (PRO) into Clinical Care
Michael Bass, Paul A. Harris, Robert J. Taylor, Justin Starren, Joshua Spuhl, Jimmy Johnson, Jason Guattery |
AMIA | 2 |
| 2016 | A multi-institution evaluation of clinical profile anonymizationabstractBACKGROUND AND OBJECTIVE: There is an increasing desire to share de-identified electronic health records (EHRs) for secondary uses, but there are concerns that clinical terms can be exploited to compromise patient identities. Anonymization algorithms mitigate such threats while enabling novel discoveries, but their evaluation has been limited to single institutions. Here, we study how an existing clinical profile anonymization fares at multiple medical centers. METHODS: We apply a state-of-the-artk-anonymization algorithm, withkset to the standard value 5, to the International Classification of Disease, ninth edition codes for patients in a hypothyroidism association study at three medical centers: Marshfield Clinic, Northwestern University, and Vanderbilt University. We assess utility when anonymizing at three population levels: all patients in 1) the EHR system; 2) the biorepository; and 3) a hypothyroidism study. We evaluate utility using 1) changes to the number included in the dataset, 2) number of codes included, and 3) regions generalization and suppression were required. RESULTS: Our findings yield several notable results. First, we show that anonymizing in the context of the entire EHR yields a significantly greater quantity of data by reducing the amount of generalized regions from ∼15% to ∼0.5%. Second, ∼70% of codes that needed generalization only generalized two or three codes in the largest anonymization. CONCLUSIONS: Sharing large volumes of clinical data in support of phenome-wide association studies is possible while safeguarding privacy to the underlying individuals. Raymond Heatherly, Luke V. Rasmussen, Peggy L. Peissig, Jennifer A. Pacheco, Paul A. Harris, Joshua C. Denny, Bradley A. Malin |
J. Am. Medical Informatics Assoc. | 5 |
| 2016 | PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportabilityabstractOBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data. Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 17 |
| 2015 | Desiderata for computable representations of electronic health records-driven phenotype algorithmsabstractBACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages. Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 29 |
| 2014 | Brief communication: The Mid-South Clinical Data Research NetworkabstractThe Mid-South Clinical Data Research Network (CDRN) encompasses three large health systems: (1) Vanderbilt Health System (VU) with electronic medical records for over 2 million patients, (2) the Vanderbilt Healthcare Affiliated Network (VHAN) which currently includes over 40 hospitals, hundreds of ambulatory practices, and over 3 million patients in the Mid-South, and (3) Greenway Medical Technologies, with access to 24 million patients nationally. Initial goals of the Mid-South CDRN include: (1) expansion of our VU data network to include the VHAN and Greenway systems, (2) developing data integration/interoperability across the three systems, (3) improving our current tools for extracting clinical data, (4) optimization of tools for collection of patient-reported data, and (5) expansion of clinical decision support. By 18 months, we anticipate our CDRN will robustly support projects in comparative effectiveness research, pragmatic clinical trials, and other key research areas and have the capacity to share data and health information technology tools nationally. S. Trent Rosenbloom, Paul A. Harris, Jill M. Pulley, Melissa A. Basford, Jason Grant, Allison DuBuisson, Russell L. Rothman |
J. Am. Medical Informatics Assoc. | 2 |
| 2014 | Secondary use of clinical data: The Vanderbilt approach
Ioana Danciu, James D. Cowan, Melissa A. Basford, Alexander Saip, Susan Osgood, Jana Shirey-Rice, Jacqueline Kirby, Paul A. Harris |
J. Biomed. Informatics | 9 |
| 2013 | Procurement of shared data instruments for Research Electronic Data Capture (REDCap)
Jihad S. Obeid, Catherine A. McGraw, Brenda L. Minor, Jose G. Conde, Robert Pawluk, Michael Lin, Janey Wang, Sean R. Banks, Sheree A. Hemphill, Robert Taylor 0001, Paul A. Harris |
J. Biomed. Informatics | 11 |
| 2012 | Research Electronic Data Capture (REDCap) - planning, collecting and managing data for clinical and translational researchabstractBackground REDCap (Research Electronic Data Capture) is a software application and workflow methodology designed to collect and manage data for research studies. REDCap study databases are secure, web-based applications and easy to create, launch and manage on a project-by-project basis. REDCap uses a study-specific data dictionary to eliminate all programming requirements for the creation of electronic case report forms and participant survey instruments for individual studies – making it extremely fast to develop and launch for any size study. Vanderbilt developed and launched REDCap in 2004 and began sharing the software with other academic and non-profit institutions in 2005 at no cost under a unique consortium dissemination model. The consortium now consists of 322 academic and non-profit partner institutions across six continents serving 38,600 end-users (http://www.projectredcap.org). This presentation will provide a description of the REDCap software platform, global consortium and low-cost institutional models for supporting data management across the entire clinical and translational research enterprise. Paul A. Harris |
BMC Bioinform. | 1 |
| 2011 | StarBRITE: The Vanderbilt University Biomedical Research Integration, Translation and Education portal
Paul A. Harris, Jonathan A. Swafford, Terri L. Edwards, Minhua Zhang, Shraddha S. Nigavekar, Tonya R. Yarbrough, Lynda D. Lane, Tara Helmer, Laurie A. Lebo, Gail Mayo, Daniel R. Masys, Gordon R. Bernard, Jill M. Pulley |
J. Biomed. Informatics | 1 |
| 2009 | Research electronic data capture (REDCap) - A metadata-driven methodology and workflow process for providing translational research informatics support
Paul A. Harris, Robert Taylor 0001, Robert Thielke, Jonathon Payne, Nathaniel Gonzalez, Jose G. Conde |
J. Biomed. Informatics | 1 |
| 2005 | Application of Information Technology: Clinical Research Subject Recruitment: The Volunteer for Vanderbilt Research Program www.volunteer.mc.vanderbilt.eduabstractThis article provides information concerning a novel research subject recruitment registry developed at Vanderbilt University. Project goals were (1) to provide a mechanism for lay individuals to self-enter information conveying interest in volunteering for clinical research and (2) provide tools for researchers to select and contact potential volunteers based on study-specific inclusion criteria. The registry was built and offered as an institutional resource to all university scientists conducting institutional review board-approved research. The authors present (1) a model for redesigning workflow associated with subject registration, volunteer retrieval, and subject contact; (2) details of a Web-based software application used as a focal point in designing workflow for our system; (3) descriptive statistics for volunteer and researcher use of the system during the first 32 months of operation; (4) cost estimates for the project; and (5) a set of recommendations for other medical centers wishing to adopt similar methodology. Paul A. Harris, Lynda D. Lane, Italo Biaggioni |
J. Am. Medical Informatics Assoc. | 1 |