EDBT 2026 Demo / reviewers in the wild / expert
Randi E. Foraker
dblp:134/1984
· DBLP profile ↗
16ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0001-9255-9394ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Comparison of rule- and large language model-based phenotype extraction from clinical notes for neurofibromatosis type 1abstractINTRODUCTION: Neurofibromatosis type 1 (NF1) is a rare genetic disorder affecting multiple organ systems with significant clinical heterogeneity. Managing individuals with NF1 is challenging due to variability in disease progression and outcomes and limited early risk assessment tools. OBJECTIVE: This study aims to develop an effective, generalizable, user-friendly clinical entity extraction pipeline for identifying NF1-related phenotypes from unstructured clinical notes to enhance research and risk-modeling efforts. We compare the benefits of rule-based natural language processing (NLP) vs large language models (LLMs) for this purpose. MATERIALS AND METHODS: Four phenotype extraction pipelines (3 LLM-based vs 1 rule-based) were developed to automatically extract selected NF1-relevant phenotypes. Subject matter experts manually reviewed clinical notes, generating a gold-standard annotation dataset for evaluation. In Phase 1, notes authored by a single NF1 physician were used to guide pipeline development and refinement. In Phase 2, notes from a second NF1 physician were used to assess pipeline generalizability, followed by further refinement to accommodate differences in physician terminology. RESULTS: With refinement, the rule-based model had higher distributions of F1 scores than the LLMs in both Phase 1 and Phase 2. However, the LLMs demonstrated better generalizability between physicians without refinement, showing lesser performance decreases (4.4%-5.1%) when transitioning from Phase 1 to Phase 2 without refinement, compared to an 8.8% decrease for the rule-based model. CONCLUSION: We highlight trade-offs between the effectiveness of rule-based NLP vs generalizability and ease of implementation of LLMs for clinical entity extraction, with implications for pipeline portability across providers and institutions. Levi Kaster, Ethan Hillis, Inez Y. Oh, Elizabeth C. Cordell, Randi E. Foraker, Albert M. Lai, Stephanie M. Morris, David H. Gutmann, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 5 |
| 2025 | Developing and sustaining inclusive language in biomedical informatics communications: an AMIA Board of Directors endorsed paper on the Inclusive Language and Context Style GuidelinesabstractOBJECTIVES: In 2023, AMIA's Inclusive Language and Context Style Guidelines (the "Guidelines") were approved by the Board of Directors and made a publicly available resource. This work began in 2021 through AMIA's DEI Task Force and subsequent DEI Committee; many members provided input, feedback, and time to create the Guidelines. In this paper, the authors provide a transparent account of the origin, development, contents, and dissemination of the Guidelines and share plans for their future development and use. MATERIALS AND METHODS: Our approach to drafting, refining, and distributing the Guidelines included consulting existing language guides, AMIA member reviews, external expert reviews, webinars, and workshops. Through an iterative approach to drafting and refining the Guidelines, the authors consulted relevant language guidelines and many experts throughout and beyond the AMIA community. RESULTS: The Inclusive Language Context Guidelines were formally approved by the AMIA Board of Directors on February 15, 2023. The Guidelines included four principles to be considered in scientific communications: Plurality, Precision, Transparency, and Destigmatization. DISCUSSION: A moment of vulnerability where an AMIA member raised concerns about the use of harmful language during a presentation resulted in the creation of a principled approach to support inclusive language within biomedical and health informatics communications. We envision that the Guidelines will support health equity by challenging dominant public narratives around health, fostering stronger interdisciplinary collaboration and critical thinking about the impact of language, and creating a more welcoming environment for the broader AMIA community. This work could not have been completed without the support of many AMIA members and other researchers in biomedical and health informatics. The Guidelines are a living document that will continue to be updated with input and feedback from the AMIA community into the future. Oliver J. Bear Don't Walk IV, Shefali Haldar, Duo Helen Wei, Hu Huang 0004, Rebecca L. Rivera, Jungwei Fan 0001, Vipina Kuttichi Keloth, Tiffany I. Leung, Pooja M. Desai, Diane M. Korngiebel, Lisa Grossman Liu, Adrienne Pichon, Vignesh Subbian, Tony Solomonides, Laura K. Wiley, Omolola Ogunyemi, Gretchen Purcell Jackson, Irene Dankwa-Mullan, Lisa Dirks, Avery Rose Everhart, Andrea G. Parker, Bradley E. Iott, Clair A. Kronk, Randi E. Foraker, Krista G. Martin, Tara Anand, Salvatore G. Volpe, Nathan Yung, Rubina F. Rizvi, Robert James Lucero, Tiffani J. Bright |
J. Am. Medical Informatics Assoc. | 24 |
| 2024 | Machine learning classification of new firearm injury encounters in the St Louis region: 2010-2020abstractOBJECTIVES: To improve firearm injury encounter classification (new vs follow-up) using machine learning (ML) and compare our ML model to other common approaches. MATERIALS AND METHODS: This retrospective study used data from the St Louis region-wide hospital-based violence intervention program data repository (2010-2020). We randomly selected 500 patients with a firearm injury diagnosis for inclusion, with 808 total firearm injury encounters split (70/30) for training and testing. We trained a least absolute shrinkage and selection operator (LASSO) regression model with the following predictors: admission type, time between firearm injury visits, number of prior firearm injury emergency department (ED) visits, encounter type (ED or other), and diagnostic codes. Our gold standard for new firearm injury encounter classification was manual chart review. We then used our test data to compare the performance of our ML model to other commonly used approaches (proxy measures of ED visits and time between firearm injury encounters, and diagnostic code encounter type designation [initial vs subsequent or sequela]). Performance metrics included area under the curve (AUC), sensitivity, and specificity with 95% confidence intervals (CIs). RESULTS: The ML model had excellent discrimination (0.92, 0.88-0.96) with high sensitivity (0.95, 0.90-0.98) and specificity (0.89, 0.81-0.95). AUC was significantly higher than time-based outcomes, sensitivity was slightly (but not significantly) lower than other approaches, and specificity was higher than all other methods. DISCUSSION: ML successfully delineated new firearm injury encounters, outperforming other approaches in ruling out encounters for follow-up. CONCLUSION: ML can be used to identify new firearm injury encounters and may be particularly useful in studies assessing re-injuries. Rachel M. Ancona, Benjamin P. Cooper, Randi E. Foraker, Taylor Kaser, Opeolu Adeoye, Kristen L. Mueller |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Electronic health record data quality assessment and tools: a systematic reviewabstractOBJECTIVE: We extended a 2013 literature review on electronic health record (EHR) data quality assessment approaches and tools to determine recent improvements or changes in EHR data quality assessment methodologies. MATERIALS AND METHODS: We completed a systematic review of PubMed articles from 2013 to April 2023 that discussed the quality assessment of EHR data. We screened and reviewed papers for the dimensions and methods defined in the original 2013 manuscript. We categorized papers as data quality outcomes of interest, tools, or opinion pieces. We abstracted and defined additional themes and methods though an iterative review process. RESULTS: We included 103 papers in the review, of which 73 were data quality outcomes of interest papers, 22 were tools, and 8 were opinion pieces. The most common dimension of data quality assessed was completeness, followed by correctness, concordance, plausibility, and currency. We abstracted conformance and bias as 2 additional dimensions of data quality and structural agreement as an additional methodology. DISCUSSION: There has been an increase in EHR data quality assessment publications since the original 2013 review. Consistent dimensions of EHR data quality continue to be assessed across applications. Despite consistent patterns of assessment, there still does not exist a standard approach for assessing EHR data quality. CONCLUSION: Guidelines are needed for EHR data quality assessment to improve the efficiency, transparency, comparability, and interoperability of data quality assessment. These guidelines must be both scalable and flexible. Automation could be helpful in generalizing this process. Abigail E. Lewis, Nicole Gray Weiskopf, Zachary B. Abrams, Randi E. Foraker, Albert M. Lai, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 4 |
| 2022 | Demonstrating an approach for evaluating synthetic geospatial and temporal epidemiologic data utility: results from analyzing >1.8 million SARS-CoV-2 tests in the United States National COVID Cohort Collaborative (N3C)abstractOBJECTIVE: This study sought to evaluate whether synthetic data derived from a national coronavirus disease 2019 (COVID-19) dataset could be used for geospatial and temporal epidemic analyses. MATERIALS AND METHODS: Using an original dataset (n = 1 854 968 severe acute respiratory syndrome coronavirus 2 tests) and its synthetic derivative, we compared key indicators of COVID-19 community spread through analysis of aggregate and zip code-level epidemic curves, patient characteristics and outcomes, distribution of tests by zip code, and indicator counts stratified by month and zip code. Similarity between the data was statistically and qualitatively evaluated. RESULTS: In general, synthetic data closely matched original data for epidemic curves, patient characteristics, and outcomes. Synthetic data suppressed labels of zip codes with few total tests (mean = 2.9 ± 2.4; max = 16 tests; 66% reduction of unique zip codes). Epidemic curves and monthly indicator counts were similar between synthetic and original data in a random sample of the most tested (top 1%; n = 171) and for all unsuppressed zip codes (n = 5819), respectively. In small sample sizes, synthetic data utility was notably decreased. DISCUSSION: Analyses on the population-level and of densely tested zip codes (which contained most of the data) were similar between original and synthetically derived datasets. Analyses of sparsely tested populations were less similar and had more data suppression. CONCLUSION: In general, synthetic data were successfully used to analyze geospatial and temporal trends. Analyses using small sample sizes or populations were limited, in part due to purposeful data label suppression-an attribute disclosure countermeasure. Users should consider data fitness for use in these cases. Jason A. Thomas, Randi E. Foraker, Noa Zamstein, Jon D. Morrow, Philip R. O. Payne, Adam B. Wilcox, Melissa A. Haendel, Christopher G. Chute, Kenneth R. Gersing, Anita Walden, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Justin Starren, Christine Suver, Chunlei Wu, Davera Gabriel, Stephanie S. Hong, Kristin Kostka, Harold P. Lehmann, Richard A. Moffitt, Michele Morris, Matvey Palchuk, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Mark M. Bissell, Marshall Clark, Andrew T. Girvin, Adam M. Lee, Robert T. Miller, Kellie M. Walters, Yooree Chae, Connor Cook, Alexandra Dest, Racquel R. Dietz, Thomas Dillon, Patricia A. Francis, Rafael Fuentes, Alexis Graves, Andrew J. Neumann, Shawn T. O'Neil, Usman Sheikh, Andréa M. Volz, Elizabeth Zampino, Christopher P. Austin, Samuel Bozzette, Mariam Deacy, Nicole Garbarini, Michael G. Kurilla, Samuel G. Michael, Joni L. Rutter, Meredith Temple-O'Connor, Katie Rebecca Bradwell, Amin Manna, Nabeel Qureshi, Mary Morrison Saltz, Julie A. McMurry, Carolyn T. Bramante, Jeremy Richard Harper, Wenndy Hernandez, Farrukh M. Koraishy, Federico Mariona, Saidulu Mattapally, Amit Saha, Satyanarayana Vedula, Yujuan Fu, Nisha Mathews, Ofer Mendelevitch |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Quantifying the Period Prevalence of Autoimmune Diseases Across the United States
Gagana Borra, Joshua M. Landman, Randi E. Foraker |
AMIA | 3 |
| 2021 | Demonstrations in Synthetic Data and the National COVID Cohort Collaborative (N3C)
Adam B. Wilcox, Randi E. Foraker, Jason A. Thomas, Jon D. Morrow, Noa Zamstein |
AMIA | 2 |
| 2021 | Predicting COVID-19 Regional Case Loads with Synthetic Data
Adam B. Wilcox, Noa Zamstein, Randi E. Foraker, Jason A. Thomas, Kenneth Wilkins, Jon D. Morrow |
AMIA | 3 |
| 2021 | Prediction of COVID-19 Case Severity Using Synthetic Data Derived from the National COVID Cohort Collaborative
Noa Zamstein, Andrew J. Neumann, Randi E. Foraker, Jason A. Thomas, Adam B. Wilcox, Jon D. Morrow |
AMIA | 3 |
| 2021 | Transformer-based Multi-target Regression on Electronic Health Records for Primordial Prevention of Cardiovascular DiseaseabstractMachine learning algorithms have been widely used to capture the static and temporal patterns within electronic health records (EHRs). While many studies focus on the (primary) prevention of diseases, primordial prevention (preventing the factors that are known to increase the risk of a disease occurring) is still widely under-investigated. In this study, we propose a multi-target regression model leveraging transformers to learn the bidirectional representations of EHR data and predict the future values of 11 major modifiable risk factors of cardiovascular disease (CVD). Inspired by the proven results of pre-training in natural language processing studies, we apply the same principles on EHR data, dividing the training of our model into two phases: pre-training and fine-tuning. We use the fine-tuned transformer model in a "multi-target regression" theme. Following this theme, we combine the 11 disjoint prediction tasks by adding shared and target-specific layers to the model and jointly train the entire model. We evaluate the performance of our proposed method on a large publicly available EHR dataset. Through various experiments, we demonstrate that the proposed method obtains a significant improvement (12.6% MAE on average across all 11 different outputs) over the baselines. Raphael Poulain, Mehak Gupta 0001, Randi E. Foraker, Rahmatollah Beheshti |
BIBM | 3 |
| 2020 | Estimating the Prevalence of Autoimmune Diseases
Joshua M. Landman, Randi E. Foraker |
AMIA | 2 |
| 2020 | When past is not a prologue: Adapting informatics practice during a pandemicabstractData and information technology are key to every aspect of our response to the current coronavirus disease 2019 (COVID-19) pandemic-including the diagnosis of patients and delivery of care, the development of predictive models of disease spread, and the management of personnel and equipment. The increasing engagement of informaticians at the forefront of these efforts has been a fundamental shift, from an academic to an operational role. However, the past history of informatics as a scientific domain and an area of applied practice provides little guidance or prologue for the incredible challenges that we are now tasked with performing. Building on our recent experiences, we present 4 critical lessons learned that have helped shape our scalable, data-driven response to COVID-19. We describe each of these lessons within the context of specific solutions and strategies we applied in addressing the challenges that we faced. Thomas George Kannampallil, Randi E. Foraker, Albert M. Lai, Keith F. Woeltje, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Mobile Health (mHealth) Interventions Used by Cancer Survivors to Improve Lifestyle Behavior: An Integrative Review
Marjorie M. Kelley, Jennifer Kue, Lynne Brophy, Andrea Peabody, Po-Yin Yen, Randi E. Foraker, Sharon Tucker |
AMIA | 6 |
| 2016 | The geographic distribution of cardiovascular health in the stroke prevention in healthcare delivery environments (SPHERE) study
Caryn Roth, Philip R. O. Payne, Rory C. Weier, Abigail B. Shoben, Erica N. Fletcher, Albert M. Lai, Marjorie M. Kelley, Jesse J. Plascak, Randi E. Foraker |
J. Biomed. Informatics | 9 |
| 2014 | Lessons Learned Bringing Public Health into the Primary Care Clinic through an EHR-based Application
Caryn Roth, Randi E. Foraker, Marcelo A. Lopetegui, Philip R. O. Payne |
AMIA | 2 |
| 2012 | Development of an Informatics-Based Predictive Model for 30-Day Readmission Customized for a Single Hospital System
Courtney Hebert, Jared Wasserman, Randi E. Foraker, Stanley Lemeshow, Hagop S. Mekhjian, Philip R. O. Payne, Peter J. Embí |
AMIA | 3 |