EDBT 2026 Demo / reviewers in the wild / expert
Nicole Gray Weiskopf
dblp:124/3380
· DBLP profile ↗
19ranked-venue papers
8as first author
8since 2021 · last 2024
0000-0003-0365-909XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 8 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prediction of multiclass surgical outcomes in glaucoma using multimodal deep learning based on free-text operative notes and structured EHR dataabstractOBJECTIVE: Surgical outcome prediction is challenging but necessary for postoperative management. Current machine learning models utilize pre- and post-op data, excluding intraoperative information in surgical notes. Current models also usually predict binary outcomes even when surgeries have multiple outcomes that require different postoperative management. This study addresses these gaps by incorporating intraoperative information into multimodal models for multiclass glaucoma surgery outcome prediction. MATERIALS AND METHODS: We developed and evaluated multimodal deep learning models for multiclass glaucoma trabeculectomy surgery outcomes using both structured EHR data and free-text operative notes. We compare those to baseline models that use structured EHR data exclusively, or neural network models that leverage only operative notes. RESULTS: The multimodal neural network had the highest performance with a macro AUROC of 0.750 and F1 score of 0.583. It outperformed the baseline machine learning model with structured EHR data alone (macro AUROC of 0.712 and F1 score of 0.486). Additionally, the multimodal model achieved the highest recall (0.692) for hypotony surgical failure, while the surgical success group had the highest precision (0.884) and F1 score (0.775). DISCUSSION: This study shows that operative notes are an important source of predictive information. The multimodal predictive model combining perioperative notes and structured pre- and post-op EHR data outperformed other models. Multiclass surgical outcome prediction can provide valuable insights for clinical decision-making. CONCLUSIONS: Our results show the potential of deep learning models to enhance clinical decision-making for postoperative management. They can be applied to other specialties to improve surgical outcome predictions. Wei-Chun Lin, Aiyin Chen, Xubo Song, Nicole Gray Weiskopf, Michael F. Chiang, Michelle R. Hribar |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Electronic health record data quality assessment and tools: a systematic reviewabstractOBJECTIVE: We extended a 2013 literature review on electronic health record (EHR) data quality assessment approaches and tools to determine recent improvements or changes in EHR data quality assessment methodologies. MATERIALS AND METHODS: We completed a systematic review of PubMed articles from 2013 to April 2023 that discussed the quality assessment of EHR data. We screened and reviewed papers for the dimensions and methods defined in the original 2013 manuscript. We categorized papers as data quality outcomes of interest, tools, or opinion pieces. We abstracted and defined additional themes and methods though an iterative review process. RESULTS: We included 103 papers in the review, of which 73 were data quality outcomes of interest papers, 22 were tools, and 8 were opinion pieces. The most common dimension of data quality assessed was completeness, followed by correctness, concordance, plausibility, and currency. We abstracted conformance and bias as 2 additional dimensions of data quality and structural agreement as an additional methodology. DISCUSSION: There has been an increase in EHR data quality assessment publications since the original 2013 review. Consistent dimensions of EHR data quality continue to be assessed across applications. Despite consistent patterns of assessment, there still does not exist a standard approach for assessing EHR data quality. CONCLUSION: Guidelines are needed for EHR data quality assessment to improve the efficiency, transparency, comparability, and interoperability of data quality assessment. These guidelines must be both scalable and flexible. Automation could be helpful in generalizing this process. Abigail E. Lewis, Nicole Gray Weiskopf, Zachary B. Abrams, Randi E. Foraker, Albert M. Lai, Philip R. O. Payne |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Healthcare utilization is a collider: an introduction to collider bias in EHR data reuseabstractOBJECTIVES: Collider bias is a common threat to internal validity in clinical research but is rarely mentioned in informatics education or literature. Conditioning on a collider, which is a variable that is the shared causal descendant of an exposure and outcome, may result in spurious associations between the exposure and outcome. Our objective is to introduce readers to collider bias and its corollaries in the retrospective analysis of electronic health record (EHR) data. TARGET AUDIENCE: Collider bias is likely to arise in the reuse of EHR data, due to data-generating mechanisms and the nature of healthcare access and utilization in the United States. Therefore, this tutorial is aimed at informaticians and other EHR data consumers without a background in epidemiological methods or causal inference. SCOPE: We focus specifically on problems that may arise from conditioning on forms of healthcare utilization, a common collider that is an implicit selection criterion when one reuses EHR data. Directed acyclic graphs (DAGs) are introduced as a tool for identifying potential sources of bias during study design and planning. References for additional resources on causal inference and DAG construction are provided. Nicole Gray Weiskopf, David A. Dorr, Christie Jackson, Harold P. Lehmann, Caroline A. Thompson |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Validating Complex Phenotypes: A Structured Approach for Dementia
David A. Dorr, Nicole Gray Weiskopf, Michelle Bobo, MJ Dunne, Peijan Han, Jessica Kim, V. G. Vinod Vydiswaran |
AMIA | 2 |
| 2022 | Comparing ascertainment of chronic condition status with problem lists versus encounter diagnoses from electronic health recordsabstractOBJECTIVE: To assess and compare electronic health record (EHR) documentation of chronic disease in problem lists and encounter diagnosis records among Community Health Center (CHC) patients. MATERIALS AND METHODS: We assessed patient EHR data in a large clinical research network during 2012-2019. We included CHCs who provided outpatient, older adult primary care to patients age ≥45 years, with ≥2 office visits during the study. Our study sample included 1 180 290 patients from 545 CHCs across 22 states. We used diagnosis codes from 39 Chronic Condition Warehouse algorithms to identify chronic conditions from encounter diagnoses only and compared against problem list records. We measured correspondence including agreement, kappa, prevalence index, bias index, and prevalence-adjusted bias-adjusted kappa. RESULTS: Overlap of encounter diagnosis and problem list ascertainment was 59.4% among chronic conditions identified, with 12.2% of conditions identified only in encounters and 28.4% identified only in problem lists. Rates of coidentification varied by condition from 7.1% to 84.4%. Greatest agreement was found in diabetes (84.4%), HIV (78.1%), and hypertension (74.7%). Sixteen conditions had <50% agreement, including cancers and substance use disorders. Overlap for mental health conditions ranged from 47.4% for anxiety to 59.8% for depression. DISCUSSION: Agreement between the 2 sources varied substantially. Conditions requiring regular management in primary care settings may have a higher agreement than those diagnosed and treated in specialty care. CONCLUSION: Relying on EHR encounter data to identify chronic conditions without reference to patient problem lists may under-capture conditions among CHC patients in the United States. Robert W. Voss, Teresa D. Schmidt, Nicole Gray Weiskopf, Miguel Marino, David A. Dorr, Nathalie Huguet, Nate Warren, Steele Valenzuela, Jean P. O'Malley, Ana R. Quiñones |
J. Am. Medical Informatics Assoc. | 3 |
| 2021 | Chart Completion Time of Attending Physicians While Using Medical Scribes
Sarah T. Florig, Sky M. Corby, Nicholas T. Rosson, Tanuj Devara, Nicole Gray Weiskopf, Jeffrey Allen Gold, Vishnu Mohan |
AMIA | 5 |
| 2021 | Extracting Patient-level Social Determinants of Health into the OMOP Common Data Model
Jimmy Phuong, Elizabeth Zampino, Nicholas J. Dobbins, Juan Espinoza, Daniella Meeker, Heidi Spratt, Charisse R. Madlock-Brown, Nicole Gray Weiskopf, Adam B. Wilcox |
AMIA | 8 |
| 2021 | The quality of social determinants data in the electronic health record: a systematic reviewabstractOBJECTIVE: The aim of this study was to collect and synthesize evidence regarding data quality problems encountered when working with variables related to social determinants of health (SDoH). MATERIALS AND METHODS: We conducted a systematic review of the literature on social determinants research and data quality and then iteratively identified themes in the literature using a content analysis process. RESULTS: The most commonly represented quality issue associated with SDoH data is plausibility (n = 31, 41%). Factors related to race and ethnicity have the largest body of literature (n = 40, 53%). The first theme, noted in 62% (n = 47) of articles, is that bias or validity issues often result from data quality problems. The most frequently identified validity issue is misclassification bias (n = 23, 30%). The second theme is that many of the articles suggest methods for mitigating the issues resulting from poor social determinants data quality. We grouped these into 5 suggestions: avoid complete case analysis, impute data, rely on multiple sources, use validated software tools, and select addresses thoughtfully. DISCUSSION: The type of data quality problem varies depending on the variable, and each problem is associated with particular forms of analytical error. Problems encountered with the quality of SDoH data are rarely distributed randomly. Data from Hispanic patients are more prone to issues with plausibility and misclassification than data from other racial/ethnic groups. CONCLUSION: Consideration of data quality and evidence-based quality improvement methods may help prevent bias and improve the validity of research conducted with SDoH data. Lily A. Cook, Jonathan Sachs, Nicole Gray Weiskopf |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | Bias in the Reuse and Analysis of Electronic Health Record Data
Nicole Gray Weiskopf, Melody L. Greer, Karthik Natarajan, Caroline A. Thompson, Harold P. Lehmann |
AMIA | 1 |
| 2019 | Towards augmenting structured EHR data: a comparison of manual chart review and patient self-report
Nicole Gray Weiskopf, Aaron M. Cohen, Joely Hannan, Thad Jarmon, David A. Dorr |
AMIA | 1 |
| 2017 | Specifications of Clinical Quality Measures and Value Set Vocabularies Shift Over Time: A Study of Change through Implementation Differences
Raja Arul Cholan, Nicole Gray Weiskopf, Doug Rhoton, Nicholas V. Colin, Rachel L. Ross, Melanie N. Marzullo, Bhavaya Sachdeva, David A. Dorr |
AMIA | 2 |
| 2017 | A Framework for Data Quality Assessment in Clinical Research Datasets
Kathleen Lee, Nicole Gray Weiskopf, Jyotishman Pathak |
AMIA | 2 |
| 2016 | Comparison of Electronic Health Record Data Sources to a Gold Standard Patient Data Set in Correctly Identifying Chronic Conditions
Shelby J. Martin, Nicole Gray Weiskopf, David A. Dorr |
AMIA | 2 |
| 2016 | A Mixed Methods Task Analysis of the Implementation and Validation of EHR-Based Clinical Quality Measures
Nicole Gray Weiskopf, Faiza Khan, David A. Dorr, Deborah V. Woodcock, Joaquin E. Cigarroa, Aaron M. Cohen |
AMIA | 1 |
| 2015 | A Guideline for Assessing EHR Data Quality for Secondary Use
Nicole Gray Weiskopf, Chunhua Weng |
AMIA | 1 |
| 2014 | Diagnosis code assignment: models and evaluation metricsabstractBACKGROUND AND OBJECTIVE: The volume of healthcare data is growing rapidly with the adoption of health information technology. We focus on automated ICD9 code assignment from discharge summary content and methods for evaluating such assignments. METHODS: We study ICD9 diagnosis codes and discharge summaries from the publicly available Multiparameter Intelligent Monitoring in Intensive Care II (MIMIC II) repository. We experiment with two coding approaches: one that treats each ICD9 code independently of each other (flat classifier), and one that leverages the hierarchical nature of ICD9 codes into its modeling (hierarchy-based classifier). We propose novel evaluation metrics, which reflect the distances among gold-standard and predicted codes and their locations in the ICD9 tree. Experimental setup, code for modeling, and evaluation scripts are made available to the research community. RESULTS: The hierarchy-based classifier outperforms the flat classifier with F-measures of 39.5% and 27.6%, respectively, when trained on 20,533 documents and tested on 2282 documents. While recall is improved at the expense of precision, our novel evaluation metrics show a more refined assessment: for instance, the hierarchy-based classifier identifies the correct sub-tree of gold-standard codes more often than the flat classifier. Error analysis reveals that gold-standard codes are not perfect, and as such the recall and precision are likely underestimated. CONCLUSIONS: Hierarchy-based classification yields better ICD9 coding than flat classification for MIMIC patients. Automated ICD9 coding is an example of a task for which data and tools can be shared and for which the research community can work together to build on shared models and advance the state of the art. Adler J. Perotte, Rimma Perotte, Karthik Natarajan, Nicole Gray Weiskopf, Frank D. Wood, Noémie Elhadad |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Sick Patients Have More Data: The Non-Random Completeness of Electronic Health Records
Nicole Gray Weiskopf, Alex Rusanov, Chunhua Weng |
AMIA | 1 |
| 2013 | Methods and dimensions of electronic health record data quality assessment: enabling reuse for clinical researchabstractOBJECTIVE: To review the methods and dimensions of data quality assessment in the context of electronic health record (EHR) data reuse for research. MATERIALS AND METHODS: A review of the clinical research literature discussing data quality assessment methodology for EHR data was performed. Using an iterative process, the aspects of data quality being measured were abstracted and categorized, as well as the methods of assessment used. RESULTS: Five dimensions of data quality were identified, which are completeness, correctness, concordance, plausibility, and currency, and seven broad categories of data quality assessment methods: comparison with gold standards, data element agreement, data source agreement, distribution comparison, validity checks, log review, and element presence. DISCUSSION: Examination of the methods by which clinical researchers have investigated the quality and suitability of EHR data for research shows that there are fundamental features of data quality, which may be difficult to measure, as well as proxy dimensions. Researchers interested in the reuse of EHR data for clinical research are recommended to consider the adoption of a consistent taxonomy of EHR data quality, to remain aware of the task-dependence of data quality, to integrate work on data quality assessment from other fields, and to adopt systematic, empirically driven, statistically based methods of data quality assessment. CONCLUSION: There is currently little consistency or potential generalizability in the methods used to assess EHR data quality. If the reuse of EHR data for clinical research is to become accepted, researchers should adopt validated, systematic methods of EHR data quality assessment. Nicole Gray Weiskopf, Chunhua Weng |
J. Am. Medical Informatics Assoc. | 1 |
| 2013 | Defining and measuring completeness of electronic health records for secondary useabstractWe demonstrate the importance of explicit definitions of electronic health record (EHR) data completeness and how different conceptualizations of completeness may impact findings from EHR-derived datasets. This study has important repercussions for researchers and clinicians engaged in the secondary use of EHR data. We describe four prototypical definitions of EHR completeness: documentation, breadth, density, and predictive completeness. Each definition dictates a different approach to the measurement of completeness. These measures were applied to representative data from NewYork-Presbyterian Hospital's clinical data warehouse. We found that according to any definition, the number of complete records in our clinical database is far lower than the nominal total. The proportion that meets criteria for completeness is heavily dependent on the definition of completeness used, and the different definitions generate different subsets of records. We conclude that the concept of completeness in EHR is contextual. We urge data consumers to be explicit in how they define a complete record and transparent about the limitations of their data. Nicole Gray Weiskopf, George Hripcsak, Sushmita Swaminathan, Chunhua Weng |
J. Biomed. Informatics | 1 |