EDBT 2026 Demo / reviewers in the wild / expert
Robert M. Cronin
dblp:133/5965
· DBLP profile ↗
29ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0003-1916-6521ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Balancing efficacy and computational burden: weighted mean, multiple imputation, and inverse probability weighting methods for item non-response in reliable scalesabstractIMPORTANCE: Scales often arise from multi-item questionnaires, yet commonly face item non-response. Traditional solutions use weighted mean (WMean) from available responses, but potentially overlook missing data intricacies. Advanced methods like multiple imputation (MI) address broader missing data, but demand increased computational resources. Researchers frequently use survey data in the All of Us Research Program (All of Us), and it is imperative to determine if the increased computational burden of employing MI to handle non-response is justifiable. OBJECTIVES: Using the 5-item Physical Activity Neighborhood Environment Scale (PANES) in All of Us, this study assessed the tradeoff between efficacy and computational demands of WMean, MI, and inverse probability weighting (IPW) when dealing with item non-response. MATERIALS AND METHODS: Synthetic missingness, allowing 1 or more item non-response, was introduced into PANES across 3 missing mechanisms and various missing percentages (10%-50%). Each scenario compared WMean of complete questions, MI, and IPW on bias, variability, coverage probability, and computation time. RESULTS: All methods showed minimal biases (all <5.5%) for good internal consistency, with WMean suffered most with poor consistency. IPW showed considerable variability with increasing missing percentage. MI required significantly more computational resources, taking >8000 and >100 times longer than WMean and IPW in full data analysis, respectively. DISCUSSION AND CONCLUSION: The marginal performance advantages of MI for item non-response in highly reliable scales do not warrant its escalated cloud computational burden in All of Us, particularly when coupled with computationally demanding post-imputation analyses. Researchers using survey scales with low missingness could utilize WMean to reduce computing burden. Andrew Guide, Shawn Garbett, Xiaoke Feng, Brandy Mapes, Justin Cook, Lina M. Sulieman, Robert M. Cronin, Qingxia Chen |
J. Am. Medical Informatics Assoc. | 7 |
| 2024 | Identifying erroneous height and weight values from adult electronic health records in the All of Us research programabstractINTRODUCTION: Electronic Health Records (EHR) are a useful data source for research, but their usability is hindered by measurement errors. This study investigated an automatic error detection algorithm for adult height and weight measurements in EHR for the All of Us Research Program (All of Us). METHODS: We developed reference charts for adult heights and weights that were stratified on participant sex. Our analysis included 4,076,534 height and 5,207,328 wt measurements from ∼ 150,000 participants. Errors were identified using modified standard deviation scores, differences from their expected values, and significant changes between consecutive measurements. We evaluated our method with chart-reviewed heights (8,092) and weights (9,039) from 250 randomly selected participants and compared it with the current cleaning algorithm in All of Us. RESULTS: The proposed algorithm classified 1.4 % of height and 1.5 % of weight errors in the full cohort. Sensitivity was 90.4 % (95 % CI: 79.0-96.8 %) for heights and 65.9 % (95 % CI: 56.9-74.1 %) for weights. Precision was 73.4 % (95 % CI: 60.9-83.7 %) for heights and 62.9 (95 % CI: 54.0-71.1 %) for weights. In comparison, the current cleaning algorithm has inferior performance in sensitivity (55.8 %) and precision (16.5 %) for height errors while having higher precision (94.0 %) and lower sensitivity (61.9 %) for weight errors. DISCUSSION: Our proposed algorithm outperformed in detecting height errors compared to weights. It can serve as a valuable addition to the current All of Us cleaning algorithm for identifying erroneous height values. Andrew Guide, Lina M. Sulieman, Shawn Garbett, Robert M. Cronin, Matthew E. Spotnitz, Karthik Natarajan, Robert J. Carroll, Paul A. Harris, Qingxia Chen |
J. Biomed. Informatics | 4 |
| 2022 | Predicting the Retention of Subsequent Surveys in the All of Us
Lina M. Sulieman, Xiaoke Feng, Qingxia Chen, Robert M. Cronin |
AMIA | 4 |
| 2022 | Comparing medical history data derived from electronic health records and survey answers in the All of Us Research ProgramabstractOBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases. Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Comparison of family health history in surveys vs electronic health record data mapped to the observational medical outcomes partnership data model in the All of Us Research ProgramabstractOBJECTIVE: Family health history is important to clinical care and precision medicine. Prior studies show gaps in data collected from patient surveys and electronic health records (EHRs). The All of Us Research Program collects family history from participants via surveys and EHRs. This Demonstration Project aims to evaluate availability of family health history information within the publicly available data from All of Us and to characterize the data from both sources. MATERIALS AND METHODS: Surveys were completed by participants on an electronic portal. EHR data was mapped to the Observational Medical Outcomes Partnership data model. We used descriptive statistics to perform exploratory analysis of the data, including evaluating a list of medically actionable genetic disorders. We performed a subanalysis on participants who had both survey and EHR data. RESULTS: There were 54 872 participants with family history data. Of those, 26% had EHR data only, 63% had survey only, and 10.5% had data from both sources. There were 35 217 participants with reported family history of a medically actionable genetic disorder (9% from EHR only, 89% from surveys, and 2% from both). In the subanalysis, we found inconsistencies between the surveys and EHRs. More details came from surveys. When both mentioned a similar disease, the source of truth was unclear. CONCLUSIONS: Compiling data from both surveys and EHR can provide a more comprehensive source for family health history, but informatics challenges and opportunities exist. Access to more complete understanding of a person's family health history may provide opportunities for precision medicine. Robert M. Cronin, Alese E. Halvorson, Cassie Springer, Xiaoke Feng, Lina M. Sulieman, Roxana Loperena-Cortes, Kelsey R. Mayo, Robert J. Carroll, Qingxia Chen, Brian K. Ahmedani, Jason Karnes, Bruce Korf, Christopher J. O'Donnell, Andrea H. Ramirez |
J. Am. Medical Informatics Assoc. | 1 |
| 2019 | Quality Analysis of the All of Us Research Program Health Surveys
Robert M. Cronin, Sarah Feng, Brandy Mapes, Roxana Loperena-Cortes, Regina Andrade, David Schlundt, Ken Wallston, Mick P. Couper, Scott Sutherland, Cindy Chen, Joshua C. Denny |
AMIA | 1 |
| 2019 | Sickle Cell Disease Phenotype Algorithm Performance in Adult, Pediatric, and Transitional Care
Mirza S. Khan, Mark Rodeghier, Robert M. Cronin |
AMIA | 3 |
| 2019 | Determinants of Medication Adherence in Sickle Cell Disease Using the World Health Organization Model
Kinsley Ojukwu, Christopher L. Simpson, Amol Utrankar, Whitney Allen, Laurie L. Novak, Robert M. Cronin |
AMIA | 6 |
| 2018 | Development of a Technology-Supported, Lay Peer-to-Peer Family Engagement Consultation Service in a Pediatric Hospital
Wayne H. Liang, Avi Madan-Swain, Robert M. Cronin, Gretchen Purcell Jackson |
AMIA | 3 |
| 2018 | The Role of Information Technologies in Sickle Cell Disease Support Systems
Amol Utrankar, Whitney Allen, Laurie L. Novak, Adetola A. Kassim, Gretchen Purcell Jackson, Michael R. DeBaun, Robert M. Cronin |
AMIA | 7 |
| 2018 | Mining 100 million notes to find homelessness and adverse childhood experiences: 2 case studies of rare and severe social determinants of health in electronic health recordsabstractObjective: Understanding how to identify the social determinants of health from electronic health records (EHRs) could provide important insights to understand health or disease outcomes. We developed a methodology to capture 2 rare and severe social determinants of health, homelessness and adverse childhood experiences (ACEs), from a large EHR repository. Materials and Methods: We first constructed lexicons to capture homelessness and ACE phenotypic profiles. We employed word2vec and lexical associations to mine homelessness-related words. Next, using relevance feedback, we refined the 2 profiles with iterative searches over 100 million notes from the Vanderbilt EHR. Seven assessors manually reviewed the top-ranked results of 2544 patient visits relevant for homelessness and 1000 patients relevant for ACE. Results: word2vec yielded better performance (area under the precision-recall curve [AUPRC] of 0.94) than lexical associations (AUPRC = 0.83) for extracting homelessness-related words. A comparative study of searches for the 2 phenotypes revealed a higher performance achieved for homelessness (AUPRC = 0.95) than ACE (AUPRC = 0.79). A temporal analysis of the homeless population showed that the majority experienced chronic homelessness. Most ACE patients suffered sexual (70%) and/or physical (50.6%) abuse, with the top-ranked abuser keywords being "father" (21.8%) and "mother" (15.4%). Top prevalent associated conditions for homeless patients were lack of housing (62.8%) and tobacco use disorder (61.5%), while for ACE patients it was mental disorders (36.6%-47.6%). Conclusion: We provide an efficient solution for mining homelessness and ACE information from EHRs, which can facilitate large clinical and genetic studies of these social determinants of health. Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | Patient and healthcare provider views on a patient-reported outcomes portalabstractBackground: Over the past decade, public interest in managing health-related information for personal understanding and self-improvement has rapidly expanded. This study explored aspects of how patient-provided health information could be obtained through an electronic portal and presented to inform and engage patients while also providing information for healthcare providers. Methods: We invited participants using ResearchMatch from 2 cohorts: (1) self-reported healthy volunteers (no medical conditions) and (2) individuals with a self-reported diagnosis of anxiety and/or depression. Participants used a secure web application (dashboard) to complete the PROMIS® domain survey(s) and then complete a feedback survey. A community engagement studio with 5 healthcare providers assessed perspectives on the feasibility and features of a portal to collect and display patient provided health information. We used bivariate analyses and regression analyses to determine differences between cohorts. Results: A total of 480 participants completed the study (239 healthy, 241 anxiety and/or depression). While participants from the tw2o cohorts had significantly different PROMIS scores (p < .05), both cohorts welcomed the concept of a patient-centric dashboard, saw value in sharing results with their healthcare provider, and wanted to view results over time. However, factors needing consideration before widespread use included personalization for the patient and their health issues, integration with existing information (eg electronic health records), and integration into clinician workflow. Conclusions: Our findings demonstrated a strong desire among healthy people, patients with chronic diseases, and healthcare providers for a self-assessment portal that can collect patient-reported outcome metrics and deliver personalized feedback. Robert M. Cronin, Douglas Conway, David M. Condon, Rebecca N. Jerome, Daniel W. Byrne, Paul A. Harris |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | A technology-based patient and family engagement consult service for the pediatric hospital settingabstractObjective: The Vanderbilt Children's Hospital launched an innovative Technology-Based Patient and Family Engagement Consult Service in 2014. This paper describes our initial experience with this service, characterizes health-related needs of families of hospitalized children, and details the technologies recommended to promote engagement and meet needs. Materials and Methods: We retrospectively reviewed consult service documentation for patient characteristics, health-related needs, and consultation team recommendations. Needs were categorized using a consumer health needs taxonomy. Recommendations were classified by technology type. Results: Twenty-two consultations were conducted with families of patients ranging in age from newborn to 15 years, most with new diagnoses or chronic illnesses. The consultation team identified 99 health-related needs (4.5 per consultation) and made 166 recommendations (7.5 per consultation, 1.7 per need). Need categories included 38 informational needs, 26 medical needs, 23 logistical needs, and 12 social needs. The most common recommendations were websites (50, 30%) and mobile applications (30, 18%). The most frequent recommendations by need category were websites for informational needs (39, 50%), mobile applications for medical needs (15, 40%), patient portals for logistical needs (12, 44%), and disease-specific support groups for social needs (19, 56%). Discussion: Families of hospitalized pediatric patients have a variety of health-related needs, many of which could be addressed by technology recommendations from an engagement consult service. Conclusion: This service is the first of its kind, offering a potentially generalizable and scalable approach to assessing health-related needs, meeting them with technologies, and promoting patient and family engagement in the inpatient setting. Gretchen Purcell Jackson, Jamie R. Robinson, Ebone Ingram, Mary Masterman, Catherine Ivory, Diane Holloway, Shilo Anders, Robert M. Cronin |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | Technology use and preferences to support clinical practice guideline awareness and adherence in individuals with sickle cell diseaseabstractObjective: Sickle cell disease (SCD) is a chronic condition affecting over 100 000 individuals in the United States, predominantly from vulnerable populations. Clinical practice guidelines, written for providers, have low adherence. This study explored knowledge about guidelines; desire for guidelines; and how technology could support guideline awareness and adherence, examining current technology uses, and user preferences to inform design of a patient-centered guidelines application in a chronic disease. Methods: This cross-sectional mixed-methods study involved semi-structured interviews, surveys, and focus groups of adolescents and adults with SCD. We evaluated interest, preferences, and anticipated benefits or barriers of a patient-centered adaptation of SCD practice guidelines; prospective technology uses for health; and barriers to technology utilization. Results: Forty-seven individuals completed surveys and interviews, and 39 participated in three separate focus groups. Most participants (91%) were unaware of SCD guidelines, but almost all (96%) expressed interest in a guidelines application, identifying benefits (knowledge, activation, individualization, and rewards), and barriers (poor information, low motivation, and resource limitations). Current technology health uses included information access, care coordination, and reminders about health-related actions. Prospective technology uses included informational messaging and timely alerts. Barriers to technology use included lack of interest, lack of utility, and preference for direct communication. Conclusions: This study's findings can inform the design of clinical practice guideline applications, suggesting a promising role for technology to engage patients, facilitate care decisions and actions, and improve outcomes. Amol Utrankar, Tilicia L. Mayo-Gamble, Whitney Allen, Laurie L. Novak, Adetola A. Kassim, Kemberlee Bonnet, David Schlundt, Velma M. Murry, Gretchen Purcell Jackson, Michael R. DeBaun, Robert M. Cronin |
J. Am. Medical Informatics Assoc. | 11 |
| 2017 | Large-Scale Text Mining of Social Determinants from Electronic Health Records: Case Studies of Homelessness and Adverse Childhood Experiences
Cosmin Adrian Bejan, John Angiolillo, Douglas Conway, Robertson Nash, Jana Shirey-Rice, Loren Lipworth-Elliot, Robert M. Cronin, Jill M. Pulley, Sunil Kripalani, Shari Barkin, Kevin B. Johnson, Joshua C. Denny |
AMIA | 7 |
| 2017 | Meeting Common Health-Related Needs Through a Pediatric Inpatient Engagement Consultation Service
Daniel J. Lee, Robert M. Cronin, Kim M. Unertl, Jamie R. Robinson, Katherine Kelly, Shilo Anders, Jennifer Wilkens, Gretchen Purcell Jackson |
AMIA | 2 |
| 2017 | Evaluating electronic health record data sources and algorithmic approaches to identify hypertensive individualsabstractOBJECTIVE: Phenotyping algorithms applied to electronic health record (EHR) data enable investigators to identify large cohorts for clinical and genomic research. Algorithm development is often iterative, depends on fallible investigator intuition, and is time- and labor-intensive. We developed and evaluated 4 types of phenotyping algorithms and categories of EHR information to identify hypertensive individuals and controls and provide a portable module for implementation at other sites. MATERIALS AND METHODS: We reviewed the EHRs of 631 individuals followed at Vanderbilt for hypertension status. We developed features and phenotyping algorithms of increasing complexity. Input categories included International Classification of Diseases, Ninth Revision (ICD9) codes, medications, vital signs, narrative-text search results, and Unified Medical Language System (UMLS) concepts extracted using natural language processing (NLP). We developed a module and tested portability by replicating 10 of the best-performing algorithms at the Marshfield Clinic. RESULTS: Random forests using billing codes, medications, vitals, and concepts had the best performance with a median area under the receiver operator characteristic curve (AUC) of 0.976. Normalized sums of all 4 categories also performed well (0.959 AUC). The best non-NLP algorithm combined normalized ICD9 codes, medications, and blood pressure readings with a median AUC of 0.948. Blood pressure cutoffs or ICD9 code counts alone had AUCs of 0.854 and 0.908, respectively. Marshfield Clinic results were similar. CONCLUSION: This work shows that billing codes or blood pressure readings alone yield good hypertension classification performance. However, even simple combinations of input categories improve performance. The most complex algorithms classified hypertension with excellent recall and precision. Pedro L. Teixeira, Wei-Qi Wei, Robert M. Cronin, Huan Mo, Jacob P. VanHouten, Robert J. Carroll, Eric LaRose, Lisa Bastarache, S. Trent Rosenbloom, Todd L. Edwards, Dan M. Roden, Thomas A. Lasko, Richard A. Dart, Anne M. Nikolai, Peggy L. Peissig, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 3 |
| 2017 | Classifying patient portal messages using Convolutional Neural Networks
Lina M. Sulieman, David Gilmore, Christi French, Robert M. Cronin, Gretchen Purcell Jackson, Matthew Russell, Daniel Fabbri |
J. Biomed. Informatics | 4 |
| 2016 | Adoption of Secure Messaging in a Patient Portal across Pediatric Specialties
Mary Masterman, Robert M. Cronin, Sharon E. Davis, Jared Shenson, Gretchen Purcell Jackson |
AMIA | 2 |
| 2016 | Use of a Patient Portal During Hospital Admissions to Surgical Services
Jamie R. Robinson, Sharon E. Davis, Robert M. Cronin, Gretchen Purcell Jackson |
AMIA | 3 |
| 2016 | Combining billing codes, clinical notes, and medications from electronic health records provides superior phenotyping performanceabstractOBJECTIVE: To evaluate the phenotyping performance of three major electronic health record (EHR) components: International Classification of Disease (ICD) diagnosis codes, primary notes, and specific medications. MATERIALS AND METHODS: We conducted the evaluation using de-identified Vanderbilt EHR data. We preselected ten diseases: atrial fibrillation, Alzheimer's disease, breast cancer, gout, human immunodeficiency virus infection, multiple sclerosis, Parkinson's disease, rheumatoid arthritis, and types 1 and 2 diabetes mellitus. For each disease, patients were classified into seven categories based on the presence of evidence in diagnosis codes, primary notes, and specific medications. Twenty-five patients per disease category (a total number of 175 patients for each disease, 1750 patients for all ten diseases) were randomly selected for manual chart review. Review results were used to estimate the positive predictive value (PPV), sensitivity, andF-score for each EHR component alone and in combination. RESULTS: The PPVs of single components were inconsistent and inadequate for accurately phenotyping (0.06-0.71). Using two or more ICD codes improved the average PPV to 0.84. We observed a more stable and higher accuracy when using at least two components (mean ± standard deviation: 0.91 ± 0.08). Primary notes offered the best sensitivity (0.77). The sensitivity of ICD codes was 0.67. Again, two or more components provided a reasonably high and stable sensitivity (0.59 ± 0.16). Overall, the best performance (Fscore: 0.70 ± 0.12) was achieved by using two or more components. Although the overall performance of using ICD codes (0.67 ± 0.14) was only slightly lower than using two or more components, its PPV (0.71 ± 0.13) is substantially worse (0.91 ± 0.08). CONCLUSION: Multiple EHR components provide a more consistent and higher performance than a single one for the selected phenotypes. We suggest considering multiple EHR components for future phenotyping design in order to obtain an ideal result. Wei-Qi Wei, Pedro L. Teixeira, Huan Mo, Robert M. Cronin, Jeremy L. Warner, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Automated Classification of Consumer Health Information Needs in Patient Portal Messages
Robert M. Cronin, Daniel Fabbri, Joshua C. Denny, Gretchen Purcell Jackson |
AMIA | 1 |
| 2015 | National Veterans Health Administration inpatient risk stratification models for hospital-acquired acute kidney injuryabstractOBJECTIVE: Hospital-acquired acute kidney injury (HA-AKI) is a potentially preventable cause of morbidity and mortality. Identifying high-risk patients prior to the onset of kidney injury is a key step towards AKI prevention. MATERIALS AND METHODS: A national retrospective cohort of 1,620,898 patient hospitalizations from 116 Veterans Affairs hospitals was assembled from electronic health record (EHR) data collected from 2003 to 2012. HA-AKI was defined at stage 1+, stage 2+, and dialysis. EHR-based predictors were identified through logistic regression, least absolute shrinkage and selection operator (lasso) regression, and random forests, and pair-wise comparisons between each were made. Calibration and discrimination metrics were calculated using 50 bootstrap iterations. In the final models, we report odds ratios, 95% confidence intervals, and importance rankings for predictor variables to evaluate their significance. RESULTS: The area under the receiver operating characteristic curve (AUC) for the different model outcomes ranged from 0.746 to 0.758 in stage 1+, 0.714 to 0.720 in stage 2+, and 0.823 to 0.825 in dialysis. Logistic regression had the best AUC in stage 1+ and dialysis. Random forests had the best AUC in stage 2+ but the least favorable calibration plots. Multiple risk factors were significant in our models, including some nonsteroidal anti-inflammatory drugs, blood pressure medications, antibiotics, and intravenous fluids given during the first 48 h of admission. CONCLUSIONS: This study demonstrated that, although all the models tested had good discrimination, performance characteristics varied between methods, and the random forests models did not calibrate as well as the lasso or logistic regression models. In addition, novel modifiable risk factors were explored and found to be significant. Robert M. Cronin, Jacob P. VanHouten, Edward D. Siew, Svetlana K. Eden, Stephan D. Fihn, Christopher D. Nielson, Josh F. Peterson, Clifton R. Baker, T. Alp Ikizler, Theodore Speroff, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | Development of a Cirrhosis-Associated Symptom/Finding Detection Tool
Jejo Koola, Robert M. Cronin, Ruth M. Reeves, Jason N. Denton, Samuel B. Ho, Michael E. Matheny |
AMIA | 2 |
| 2014 | Evaluation of Diagnosis Codes, Clinical Notes, and Medications on Identifying Subjects with a Specific Disease Phenotype
Wei-Qi Wei, Pedro L. Teixeira, Huan Mo, Robert M. Cronin, Jeremy L. Warner, Joshua C. Denny |
AMIA | 4 |
| 2014 | Research and applications: Assisted annotation of medical free text using RapTATabstractOBJECTIVE: To determine whether assisted annotation using interactive training can reduce the time required to annotate a clinical document corpus without introducing bias. MATERIALS AND METHODS: A tool, RapTAT, was designed to assist annotation by iteratively pre-annotating probable phrases of interest within a document, presenting the annotations to a reviewer for correction, and then using the corrected annotations for further machine learning-based training before pre-annotating subsequent documents. Annotators reviewed 404 clinical notes either manually or using RapTAT assistance for concepts related to quality of care during heart failure treatment. Notes were divided into 20 batches of 19-21 documents for iterative annotation and training. RESULTS: The number of correct RapTAT pre-annotations increased significantly and annotation time per batch decreased by ~50% over the course of annotation. Annotation rate increased from batch to batch for assisted but not manual reviewers. Pre-annotation F-measure increased from 0.5 to 0.6 to >0.80 (relative to both assisted reviewer and reference annotations) over the first three batches and more slowly thereafter. Overall inter-annotator agreement was significantly higher between RapTAT-assisted reviewers (0.89) than between manual reviewers (0.85). DISCUSSION: The tool reduced workload by decreasing the number of annotations needing to be added and helping reviewers to annotate at an increased rate. Agreement between the pre-annotations and reference standard, and agreement between the pre-annotations and assisted annotations, were similar throughout the annotation process, which suggests that pre-annotation did not introduce bias. CONCLUSIONS: Pre-annotations generated by a tool capable of interactive training can reduce the time required to create an annotated document corpus by up to 50%. Glenn T. Gobbel, Jennifer H. Garvin, Ruth M. Reeves, Robert M. Cronin, Julia Heavirland, Jenifer Williams, Allison Weaver, Shrimalini Jayaramaraja, Dario A. Giuse, Theodore Speroff, Steven H. Brown, Hua Xu 0001, Michael E. Matheny |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Use and Evaluation of RapTAT-Assisted Annotation for Extraction of Acute Kidney Injury Concepts from Clinical Free Text
Glenn T. Gobbel, Ruth M. Reeves, Fern FitzHenry, Diane Montella, Robert M. Cronin, Shrimalini Jayaramaraja, Theodore Speroff, Steven H. Brown, Dario A. Giuse, Michael E. Matheny |
AMIA | 5 |
| 2013 | Methods for Detection of Kidney Disease Risk Factors in Clinical Reports
Ruth M. Reeves, Glenn T. Gobbel, Shrimalini Jayaramaraja, Fern FitzHenry, Robert M. Cronin, Theodore Speroff, Michael E. Matheny |
AMIA | 5 |
| 2013 | Development and evaluation of an ensemble resource linking medications to their indicationsabstractOBJECTIVE: To create a computable MEDication Indication resource (MEDI) to support primary and secondary use of electronic medical records (EMRs). MATERIALS AND METHODS: We processed four public medication resources, RxNorm, Side Effect Resource (SIDER) 2, MedlinePlus, and Wikipedia, to create MEDI. We applied natural language processing and ontology relationships to extract indications for prescribable, single-ingredient medication concepts and all ingredient concepts as defined by RxNorm. Indications were coded as Unified Medical Language System (UMLS) concepts and International Classification of Diseases, 9th edition (ICD9) codes. A total of 689 extracted indications were randomly selected for manual review for accuracy using dual-physician review. We identified a subset of medication-indication pairs that optimizes recall while maintaining high precision. RESULTS: MEDI contains 3112 medications and 63 343 medication-indication pairs. Wikipedia was the largest resource, with 2608 medications and 34 911 pairs. For each resource, estimated precision and recall, respectively, were 94% and 20% for RxNorm, 75% and 33% for MedlinePlus, 67% and 31% for SIDER 2, and 56% and 51% for Wikipedia. The MEDI high-precision subset (MEDI-HPS) includes indications found within either RxNorm or at least two of the three other resources. MEDI-HPS contains 13 304 unique indication pairs regarding 2136 medications. The mean±SD number of indications for each medication in MEDI-HPS is 6.22 ± 6.09. The estimated precision of MEDI-HPS is 92%. CONCLUSIONS: MEDI is a publicly available, computable resource that links medications with their indications as represented by concepts and billing codes. MEDI may benefit clinical EMR applications and reuse of EMR data for research. Wei-Qi Wei, Robert M. Cronin, Hua Xu 0001, Thomas A. Lasko, Lisa Bastarache, Joshua C. Denny |
J. Am. Medical Informatics Assoc. | 2 |