David Carrell

dblp:46/8749 · also David S. Carrell · DBLP profile ↗
← Back
33ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0002-8471-0928ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 32 · 7 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Statistical methods to harmonize electronic health record data across healthcare systems: case study and lessons learned
abstract
MOTIVATION: Although common data models for electronic health record (EHR) data can facilitate multi-site data organization and querying, the same medical event may still be coded differently between healthcare systems. In this paper, we present statistical methods to identify and mitigate coding discrepancies using summary-level data, and demonstrate these methods using data from two FDA Sentinel data partners: Kaiser Permanente Washington and Kaiser Permanente Northwest. RESULTS: We first characterize differences in coding patterns, then compute a code mapping matrix to harmonize data between systems. Our findings reveal significant heterogeneity in coded EHR data, even after adopting a common data model with the same coding system, highlighting the importance of data harmonization before downstream analyses. Our study also demonstrates the effectiveness of the data harmonization approaches, which provide a foundational data quality step to promote semantic interoperability, enhance data integration, and improve the integrity of study conclusions. AVAILABILITY AND IMPLEMENTATION: Computation prototypes, including R/Python codes and examples, are included in Section 7, available as supplementary data at Bioinformatics online and will be posted on GitHub upon publication.
Yuqi Zhai, Xianshi Yu, Brian L. Hazlehurst, Denis B. Nyongesa, Daniel S. Sapp, Brian D. Williamson, David Carrell, Luesa Healy, Kara L. Cushing-Haugen, Jenna Wong, Shirley V. Wang, James S. Floyd, Kathleen Shattuck, Samuel McGown, Sarah Alam, José J. Hernández-Muñoz, Danijela Stojanovic, Sudha R. Raman, Sharon E. Davis, Tianxi Cai, Jennifer C. Nelson, Patrick J. Heagerty
Bioinform.9
2024 A general framework for developing computable clinical phenotype algorithms
abstract
OBJECTIVE: To present a general framework providing high-level guidance to developers of computable algorithms for identifying patients with specific clinical conditions (phenotypes) through a variety of approaches, including but not limited to machine learning and natural language processing methods to incorporate rich electronic health record data. MATERIALS AND METHODS: Drawing on extensive prior phenotyping experiences and insights derived from 3 algorithm development projects conducted specifically for this purpose, our team with expertise in clinical medicine, statistics, informatics, pharmacoepidemiology, and healthcare data science methods conceptualized stages of development and corresponding sets of principles, strategies, and practical guidelines for improving the algorithm development process. RESULTS: We propose 5 stages of algorithm development and corresponding principles, strategies, and guidelines: (1) assessing fitness-for-purpose, (2) creating gold standard data, (3) feature engineering, (4) model development, and (5) model evaluation. DISCUSSION AND CONCLUSION: This framework is intended to provide practical guidance and serve as a basis for future elaboration and extension.
David Carrell, James S. Floyd, Susan Gruber, Brian L. Hazlehurst, Patrick J. Heagerty, Jennifer C. Nelson, Brian D. Williamson, Robert Ball
J. Am. Medical Informatics Assoc.1
2024 Data-driven automated classification algorithms for acute health conditions: applying PheNorm to COVID-19 disease
abstract
OBJECTIVES: Automated phenotyping algorithms can reduce development time and operator dependence compared to manually developed algorithms. One such approach, PheNorm, has performed well for identifying chronic health conditions, but its performance for acute conditions is largely unknown. Herein, we implement and evaluate PheNorm applied to symptomatic COVID-19 disease to investigate its potential feasibility for rapid phenotyping of acute health conditions. MATERIALS AND METHODS: PheNorm is a general-purpose automated approach to creating computable phenotype algorithms based on natural language processing, machine learning, and (low cost) silver-standard training labels. We applied PheNorm to cohorts of potential COVID-19 patients from 2 institutions and used gold-standard manual chart review data to investigate the impact on performance of alternative feature engineering options and implementing externally trained models without local retraining. RESULTS: Models at each institution achieved AUC, sensitivity, and positive predictive value of 0.853, 0.879, 0.851 and 0.804, 0.976, and 0.885, respectively, at quantiles of model-predicted risk that maximize F1. We report performance metrics for all combinations of silver labels, feature engineering options, and models trained internally versus externally. DISCUSSION: Phenotyping algorithms developed using PheNorm performed well at both institutions. Performance varied with different silver-standard labels and feature engineering options. Models developed locally at one site also worked well when implemented externally at the other site. CONCLUSION: PheNorm models successfully identified an acute health condition, symptomatic COVID-19. The simplicity of the PheNorm approach allows it to be applied at multiple study sites with substantially reduced overhead compared to traditional approaches.
Joshua C. Smith, Brian D. Williamson, David J. Cronkite, Daniel Park, Jill M. Whitaker, Michael F. McLemore, Joshua Osmanski, Robert Winter 0003, Arvind Ramaprasan, Ann Kelley, Mary Shea, Saranrat Wittayanukorn, Danijela Stojanovic, Yueqin Zhao, Sengwee Toh, Kevin B. Johnson, David Aronoff, David Carrell
J. Am. Medical Informatics Assoc.18
2023 Characterizing variability of electronic health record-driven phenotype definitions
abstract
OBJECTIVE: The aim of this study was to analyze a publicly available sample of rule-based phenotype definitions to characterize and evaluate the variability of logical constructs used. MATERIALS AND METHODS: A sample of 33 preexisting phenotype definitions used in research that are represented using Fast Healthcare Interoperability Resources and Clinical Quality Language (CQL) was analyzed using automated analysis of the computable representation of the CQL libraries. RESULTS: Most of the phenotype definitions include narrative descriptions and flowcharts, while few provide pseudocode or executable artifacts. Most use 4 or fewer medical terminologies. The number of codes used ranges from 5 to 6865, and value sets from 1 to 19. We found that the most common expressions used were literal, data, and logical expressions. Aggregate and arithmetic expressions are the least common. Expression depth ranges from 4 to 27. DISCUSSION: Despite the range of conditions, we found that all of the phenotype definitions consisted of logical criteria, representing both clinical and operational logic, and tabular data, consisting of codes from standard terminologies and keywords for natural language processing. The total number and variety of expressions are low, which may be to simplify implementation, or authors may limit complexity due to data availability constraints. CONCLUSIONS: The phenotype definitions analyzed show significant variation in specific logical, arithmetic, and other operators but are all composed of the same high-level components, namely tabular data and logical expressions. A standard representation for phenotype definitions should support these formats and be modular to support localization and shared logic.
Pascal S. Brandt, Abel N. Kho, Yuan Luo 0001, Jennifer A. Pacheco, Theresa Walunas, Hakon Hakonarson, George Hripcsak, Cong Liu 0020, Ning Shang 0004, Chunhua Weng, Nephi Walton, David Carrell, Paul K. Crane, Eric B. Larson, Christopher G. Chute, Iftikhar J. Kullo, Robert J. Carroll, Joshua C. Denny, Andrea H. Ramirez, Wei-Qi Wei, Jyotishman Pathak, Laura K. Wiley, Rachel L. Richesson, Justin Starren, Luke V. Rasmussen
J. Am. Medical Informatics Assoc.12
2022 Data-driven automated classification algorithms for acute health conditions: Applying PheNorm to COVID-19 disease
Joshua C. Smith, Daniel Park, Jill Whitaker Bey, Michael F. McLemore, Elizabeth Hanchrow, Dax Westerman, Joshua Osmanski, Robert Winter 0003, Arvind Ramaprasan, Ann Kelley, Mary Shea, David J. Cronkite, Saranrat Wittayanukorn, Danijela Stojanovic, Yueqin Zhao, Darren Toh, Kevin B. Johnson, David Aronoff, David Carrell
AMIA19
2021 Evaluation of the Portability of Natural Language Processing-based Computable Phenotypes in the eMERGE Network
Jennifer A. Pacheco, Luke V. Rasmussen, Ken Wiley, Thomas N. Person, David J. Cronkite, Sunghwan Sohn, Shawn N. Murphy, Justin H. Gundelach, Vivian S. Gainer, Victor M. Castro, Cong Liu 0020, Todd Lingren, Frank D. Mentch, Agnes S. Sundaresan, Garrett Eickelberg, Valerie Willis, Al'ona Furmanchuk, Roshan Patel, David Carrell, Marc S. Williams, Elizabeth W. Karlson, Jodell E. Linder, Yuan Luo 0001, Chunhua Weng, Wei-Qi Wei
AMIA19
2020 Resilience of clinical text de-identified with "hiding in plain sight" to hostile reidentification attacks by human readers
abstract
OBJECTIVE: Effective, scalable de-identification of personally identifying information (PII) for information-rich clinical text is critical to support secondary use, but no method is 100% effective. The hiding-in-plain-sight (HIPS) approach attempts to solve this "residual PII problem." HIPS replaces PII tagged by a de-identification system with realistic but fictitious (resynthesized) content, making it harder to detect remaining unredacted PII. MATERIALS AND METHODS: Using 2000 representative clinical documents from 2 healthcare settings (4000 total), we used a novel method to generate 2 de-identified 100-document corpora (200 documents total) in which PII tagged by a typical automated machine-learned tagger was replaced by HIPS-resynthesized content. Four readers conducted aggressive reidentification attacks to isolate leaked PII: 2 readers from within the originating institution and 2 external readers. RESULTS: Overall, mean recall of leaked PII was 26.8% and mean precision was 37.2%. Mean recall was 9% (mean precision = 37%) for patient ages, 32% (mean precision = 26%) for dates, 25% (mean precision = 37%) for doctor names, 45% (mean precision = 55%) for organization names, and 23% (mean precision = 57%) for patient names. Recall was 32% (precision = 40%) for internal and 22% (precision =33%) for external readers. DISCUSSION AND CONCLUSIONS: Approximately 70% of leaked PII "hiding" in a corpus de-identified with HIPS resynthesis is resilient to detection by human readers in a realistic, aggressive reidentification attack scenario-more than double the rate reported in previous studies but less than the rate reported for an attack assisted by machine learning methods.
David Carrell, Bradley A. Malin, David J. Cronkite, John S. Aberdeen, Cheryl Clark, Muqun Li, Dikshya Bastakoty, Steve Nyemba, Lynette Hirschman
J. Am. Medical Informatics Assoc.1
2019 Facilitating Self-reflection about Values and Self-care Among Individuals with Chronic Conditions
abstract
Individuals with multiple chronic conditions (MCC) experience the overwhelming burden of treating MCC and frequently disagree with their providers on priorities for care. Aligning self-care with patients' values may improve healthcare for these patients. However, patients' values are not routinely discussed in clinical conversations and patients may not actively share this information with providers. In a qualitative field study, we interviewed 15 patients in their homes to investigate techniques that encourage patients to articulate values, self-care, and how they relate. Study activities facilitated self-reflection on values and self-care and produced varying responses, including: raising consciousness, evolving perspectives, identifying misalignments, and considering changes. We discuss how our findings extend prior work on supporting reflection in HCI and inform the design of tools for improving care for people with MCC.
Catherine Lim, Andrew B. L. Berry, Andrea L. Hartzler, Tad Hirsch, David Carrell, Zoë A. Bermet, James D. Ralston
CHI5
2019 The machine giveth and the machine taketh away: a parrot attack on clinical text deidentified with hiding in plain sight
abstract
OBJECTIVE: Clinical corpora can be deidentified using a combination of machine-learned automated taggers and hiding in plain sight (HIPS) resynthesis. The latter replaces detected personally identifiable information (PII) with random surrogates, allowing leaked PII to blend in or "hide in plain sight." We evaluated the extent to which a malicious attacker could expose leaked PII in such a corpus. MATERIALS AND METHODS: We modeled a scenario where an institution (the defender) externally shared an 800-note corpus of actual outpatient clinical encounter notes from a large, integrated health care delivery system in Washington State. These notes were deidentified by a machine-learned PII tagger and HIPS resynthesis. A malicious attacker obtained and performed a parrot attack intending to expose leaked PII in this corpus. Specifically, the attacker mimicked the defender's process by manually annotating all PII-like content in half of the released corpus, training a PII tagger on these data, and using the trained model to tag the remaining encounter notes. The attacker hypothesized that untagged identifiers would be leaked PII, discoverable by manual review. We evaluated the attacker's success using measures of leak-detection rate and accuracy. RESULTS: The attacker correctly hypothesized that 211 (68%) of 310 actual PII leaks in the corpus were leaks, and wrongly hypothesized that 191 resynthesized PII instances were also leaks. One-third of actual leaks remained undetected. DISCUSSION AND CONCLUSION: A malicious parrot attack to reveal leaked PII in clinical text deidentified by machine-learned HIPS resynthesis can attenuate but not eliminate the protective effect of HIPS deidentification.
David Carrell, David J. Cronkite, Muqun Li, Steve Nyemba, Bradley A. Malin, John S. Aberdeen, Lynette Hirschman
J. Am. Medical Informatics Assoc.1
2019 Enrichment sampling for a multi-site patient survey using electronic health records and census data
abstract
Objective: We describe a stratified sampling design that combines electronic health records (EHRs) and United States Census (USC) data to construct the sampling frame and an algorithm to enrich the sample with individuals belonging to rarer strata. Materials and Methods: This design was developed for a multi-site survey that sought to examine patient concerns about and barriers to participating in research studies, especially among under-studied populations (eg, minorities, low educational attainment). We defined sampling strata by cross-tabulating several socio-demographic variables obtained from EHR and augmented with census-block-level USC data. We oversampled rarer and historically underrepresented subpopulations. Results: The sampling strategy, which included USC-supplemented EHR data, led to a far more diverse sample than would have been expected under random sampling (eg, 3-, 8-, 7-, and 12-fold increase in African Americans, Asians, Hispanics and those with less than a high school degree, respectively). We observed that our EHR data tended to misclassify minority races more often than majority races, and that non-majority races, Latino ethnicity, younger adult age, lower education, and urban/suburban living were each associated with lower response rates to the mailed surveys. Discussion: We observed substantial enrichment from rarer subpopulations. The magnitude of the enrichment depends on the accuracy of the variables that define the sampling strata and the overall response rate. Conclusion: EHR and USC data may be used to define sampling strata that in turn may be used to enrich the final study sample. This design may be of particular interest for studies of rarer and understudied populations.
Nathaniel D. Mercaldo, Kyle B. Brothers, David Carrell, Ellen Wright Clayton, John J. Connolly, Ingrid A. Holm, Carol R. Horowitz, Gail P. Jarvik, Terrie E. Kitchner, Rongling Li, Catherine A. McCarty, Jennifer B. McCormick, Valerie D. McManus, Melanie F. Myers, Joshua J. Pankratz, Martha J. Shrubsole, Maureen E. Smith, Sarah C. Stallings, Janet L. Williams, Jonathan S. Schildcrout
J. Am. Medical Informatics Assoc.3
2019 Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng
J. Biomed. Informatics8
2019 Making work visible for electronic phenotype implementation: Lessons learned from the eMERGE network
Ning Shang 0004, Cong Liu 0020, Luke V. Rasmussen, Casey N. Ta, Robert J. Carroll, Barbara Benoit, Todd Lingren, Ozan Dikilitas, Frank D. Mentch, David Carrell, Wei-Qi Wei, Yuan Luo 0001, Vivian S. Gainer, Iftikhar J. Kullo, Jennifer A. Pacheco, Hakon Hakonarson, Theresa Walunas, Joshua C. Denny, Chunhua Weng
J. Biomed. Informatics10
2018 Detecting the Presence of an Individual in Phenotypic Summary Data
Yongtai Liu, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Abel N. Kho, David Carrell, Bradley A. Malin
AMIA8
2018 Automatic Text De-Identification: How and When is it Acceptable?
Stéphane M. Meystre, David Carrell, Lynette Hirschman, John S. Aberdeen, Paul Fearn, Valentina Petkov, Jonathan C. Silverstein
AMIA2
2017 Challenges in adapting existing clinical natural language processing systems to multiple, diverse health care settings
abstract
OBJECTIVE: Widespread application of clinical natural language processing (NLP) systems requires taking existing NLP systems and adapting them to diverse and heterogeneous settings. We describe the challenges faced and lessons learned in adapting an existing NLP system for measuring colonoscopy quality. MATERIALS AND METHODS: Colonoscopy and pathology reports from 4 settings during 2013-2015, varying by geographic location, practice type, compensation structure, and electronic health record. RESULTS: Though successful, adaptation required considerably more time and effort than anticipated. Typical NLP challenges in assembling corpora, diverse report structures, and idiosyncratic linguistic content were greatly magnified. DISCUSSION: Strategies for addressing adaptation challenges include assessing site-specific diversity, setting realistic timelines, leveraging local electronic health record expertise, and undertaking extensive iterative development. More research is needed on how to make it easier to adapt NLP systems to new clinical settings. CONCLUSIONS: A key challenge in widespread application of NLP is adapting existing systems to new clinical settings.
David Carrell, Robert E. Schoen, Daniel A. Leffler, Michele Morris, Sherri Rose, Andrew Baer, Seth D. Crockett, Rebecca Gourevitch, Katie M. Dean, Ateev Mehrotra
J. Am. Medical Informatics Assoc.1
2016 PheKB: a catalog and workflow for creating electronic phenotype algorithms for transportability
abstract
OBJECTIVE: Health care generated data have become an important source for clinical and genomic research. Often, investigators create and iteratively refine phenotype algorithms to achieve high positive predictive values (PPVs) or sensitivity, thereby identifying valid cases and controls. These algorithms achieve the greatest utility when validated and shared by multiple health care systems.Materials and Methods We report the current status and impact of the Phenotype KnowledgeBase (PheKB, http://phekb.org), an online environment supporting the workflow of building, sharing, and validating electronic phenotype algorithms. We analyze the most frequent components used in algorithms and their performance at authoring institutions and secondary implementation sites. RESULTS: As of June 2015, PheKB contained 30 finalized phenotype algorithms and 62 algorithms in development spanning a range of traits and diseases. Phenotypes have had over 3500 unique views in a 6-month period and have been reused by other institutions. International Classification of Disease codes were the most frequently used component, followed by medications and natural language processing. Among algorithms with published performance data, the median PPV was nearly identical when evaluated at the authoring institutions (n = 44; case 96.0%, control 100%) compared to implementation sites (n = 40; case 97.5%, control 100%). DISCUSSION: These results demonstrate that a broad range of algorithms to mine electronic health record data from different health systems can be developed with high PPV, and algorithms developed at one site are generally transportable to others. CONCLUSION: By providing a central repository, PheKB enables improved development, transportability, and validity of algorithms for research-grade phenotypes using health care generated data.
Jacqueline Kirby, Peter Speltz, Luke V. Rasmussen, Melissa A. Basford, Omri Gottesman, Peggy L. Peissig, Jennifer A. Pacheco, Gerard Tromp, Jyotishman Pathak, David Carrell, Stephen B. Ellis, Todd Lingren, William K. Thompson, Guergana K. Savova, Jonathan L. Haines, Dan M. Roden, Paul A. Harris, Joshua C. Denny
J. Am. Medical Informatics Assoc.10
2016 Optimizing annotation resources for natural language de-identification via a game theoretic framework
Muqun Li, David Carrell, John S. Aberdeen, Lynette Hirschman, Jacqueline Kirby, Bo Li 0026, Yevgeniy Vorobeychik, Bradley A. Malin
J. Biomed. Informatics2
2015 Desiderata for computable representations of electronic health records-driven phenotype algorithms
abstract
BACKGROUND: Electronic health records (EHRs) are increasingly used for clinical and translational research through the creation of phenotype algorithms. Currently, phenotype algorithms are most commonly represented as noncomputable descriptive documents and knowledge artifacts that detail the protocols for querying diagnoses, symptoms, procedures, medications, and/or text-driven medical concepts, and are primarily meant for human comprehension. We present desiderata for developing a computable phenotype representation model (PheRM). METHODS: A team of clinicians and informaticians reviewed common features for multisite phenotype algorithms published in PheKB.org and existing phenotype representation platforms. We also evaluated well-known diagnostic criteria and clinical decision-making guidelines to encompass a broader category of algorithms. RESULTS: We propose 10 desired characteristics for a flexible, computable PheRM: (1) structure clinical data into queryable forms; (2) recommend use of a common data model, but also support customization for the variability and availability of EHR data among sites; (3) support both human-readable and computable representations of phenotype algorithms; (4) implement set operations and relational algebra for modeling phenotype algorithms; (5) represent phenotype criteria with structured rules; (6) support defining temporal relations between events; (7) use standardized terminologies and ontologies, and facilitate reuse of value sets; (8) define representations for text searching and natural language processing; (9) provide interfaces for external software algorithms; and (10) maintain backward compatibility. CONCLUSION: A computable PheRM is needed for true phenotype portability and reliability across different EHR products and healthcare systems. These desiderata are a guide to inform the establishment and evolution of EHR phenotype algorithm authoring platforms and languages.
Huan Mo, William K. Thompson, Luke V. Rasmussen, Jennifer A. Pacheco, Guoqian Jiang, Richard C. Kiefer, Qian Zhu 0003, Jie Xu 0011, Enid N. H. Montague, David Carrell, Todd Lingren, Frank D. Mentch, Yizhao Ni, Firas H. Wehbe, Peggy L. Peissig, Gerard Tromp, Eric B. Larson, Christopher G. Chute, Jyotishman Pathak, Joshua C. Denny, Peter Speltz, Abel N. Kho, Gail P. Jarvik, Cosmin Adrian Bejan, Marc S. Williams, Kenneth Borthwick, Terrie E. Kitchner, Dan M. Roden, Paul A. Harris
J. Am. Medical Informatics Assoc.10
2015 Using natural language processing to extract mammographic findings
Hongyuan Gao, Erin J. Aiello Bowles, David Carrell, Diana S. M. Buist
J. Biomed. Informatics3
2014 Linking Adenomas Between Colonoscopy And Pathology Notes For PROSPR
Scott R. Halgrim, Leslie Sizemore, Edward Pham, Gabrielle Gundersen, David Carrell, Karen Wernli, Jessica Chubak, Carolyn M. Rutter
AMIA5
2014 Toward personalizing treatment for depression: predicting diagnosis and severity
abstract
OBJECTIVE: Depression is a prevalent disorder difficult to diagnose and treat. In particular, depressed patients exhibit largely unpredictable responses to treatment. Toward the goal of personalizing treatment for depression, we develop and evaluate computational models that use electronic health record (EHR) data for predicting the diagnosis and severity of depression, and response to treatment. MATERIALS AND METHODS: We develop regression-based models for predicting depression, its severity, and response to treatment from EHR data, using structured diagnosis and medication codes as well as free-text clinical reports. We used two datasets: 35,000 patients (5000 depressed) from the Palo Alto Medical Foundation and 5651 patients treated for depression from the Group Health Research Institute. RESULTS: Our models are able to predict a future diagnosis of depression up to 12 months in advance (area under the receiver operating characteristic curve (AUC) 0.70-0.80). We can differentiate patients with severe baseline depression from those with minimal or mild baseline depression (AUC 0.72). Baseline depression severity was the strongest predictor of treatment response for medication and psychotherapy. CONCLUSIONS: It is possible to use EHR data to predict a diagnosis of depression up to 12 months in advance and to differentiate between extreme baseline levels of depression. The models use commonly available data on diagnosis, medication, and clinical progress notes, making them easily portable. The ability to automatically determine severity can facilitate assembly of large patient cohorts with similar severity from multiple sites, which may enable elucidation of the moderators of treatment response in the future.
Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah
J. Am. Medical Informatics Assoc.5
2014 Design patterns for the development of electronic health record-driven phenotype extraction algorithms
Luke V. Rasmussen, William K. Thompson, Jennifer A. Pacheco, Abel N. Kho, David Carrell, Jyotishman Pathak, Peggy L. Peissig, Gerard Tromp, Joshua C. Denny, Justin Starren
J. Biomed. Informatics5
2013 Predictive Models in Mental Health: From Diagnosis to Treatment
Sandy Huang, Paea LePendu, Srinivasan Iyer 0002, Ming Tai-Seale, David Carrell, Nigam H. Shah
AMIA5
2013 Negation's Not Solved: Reconsidering Negation Annotation and Evaluation
Stephen T. Wu, Timothy A. Miller, James J. Masanz, Matthew Coarr, David Carrell, Scott R. Halgrim, David Harris 0004, Cheryl Clark
AMIA5
2013 Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text
abstract
OBJECTIVE: Secondary use of clinical text is impeded by a lack of highly effective, low-cost de-identification methods. Both, manual and automated methods for removing protected health information, are known to leave behind residual identifiers. The authors propose a novel approach for addressing the residual identifier problem based on the theory of Hiding In Plain Sight (HIPS). MATERIALS AND METHODS: HIPS relies on obfuscation to conceal residual identifiers. According to this theory, replacing the detected identifiers with realistic but synthetic surrogates should collectively render the few 'leaked' identifiers difficult to distinguish from the synthetic surrogates. The authors conducted a pilot study to test this theory on clinical narrative, de-identified by an automated system. Test corpora included 31 oncology and 50 family practice progress notes read by two trained chart abstractors and an informaticist. RESULTS: Experimental results suggest approximately 90% of residual identifiers can be effectively concealed by the HIPS approach in text containing average and high densities of personal identifying information. DISCUSSION: This pilot test suggests HIPS is feasible, but requires further evaluation. The results need to be replicated on larger corpora of diverse origin under a range of detection scenarios. Error analyses also suggest areas where surrogate generation techniques can be refined to improve efficacy. CONCLUSIONS: If these results generalize to existing high-performing de-identification systems with recall rates of 94-98%, HIPS could increase the effective de-identification rates of these systems to levels above 99% without further advancements in system recall. Additional and more rigorous assessment of the HIPS approach is warranted.
David Carrell, Bradley A. Malin, John S. Aberdeen, Samuel Bayer, Cheryl Clark, Ben Wellner, Lynette Hirschman
J. Am. Medical Informatics Assoc.1
2012 Using Electronic Health Records to Identify Heart Failure Cohorts with Differentiation for Preserved and Reduced Ejection Fraction
Suzette J. Bielinski, Jyotishman Pathak, Sunghwan Sohn, Gail P. Jarvik, David Carrell, Naveen Pereira, Véronique L. Roger
AMIA6
2012 Using PheWAS to Assess Pleiotropy of Genetic Risk Scores for Rheumatoid Arthritis and Coronary Artery Disease in the eMERGE Network
Robert J. Carroll, Katherine P. Liao, Anne E. Eyler, Lisa Bastarache, Dana C. Crawford, Peggy L. Peissig, Jyotishman Pathak, David Carrell, Abel N. Kho, Rongling Li, Daniel R. Masys, Gail P. Jarvik, Christopher G. Chute, Rex L. Chisholm, Eric B. Larson, Catherine A. McCarty, Iftikhar J. Kullo
AMIA8
2012 Importance of multi-modal approaches to effectively identify cataract cases from electronic health records
abstract
OBJECTIVE: There is increasing interest in using electronic health records (EHRs) to identify subjects for genomic association studies, due in part to the availability of large amounts of clinical data and the expected cost efficiencies of subject identification. We describe the construction and validation of an EHR-based algorithm to identify subjects with age-related cataracts. MATERIALS AND METHODS: We used a multi-modal strategy consisting of structured database querying, natural language processing on free-text documents, and optical character recognition on scanned clinical images to identify cataract subjects and related cataract attributes. Extensive validation on 3657 subjects compared the multi-modal results to manual chart review. The algorithm was also implemented at participating electronic MEdical Records and GEnomics (eMERGE) institutions. RESULTS: An EHR-based cataract phenotyping algorithm was successfully developed and validated, resulting in positive predictive values (PPVs) >95%. The multi-modal approach increased the identification of cataract subject attributes by a factor of three compared to single-mode approaches while maintaining high PPV. Components of the cataract algorithm were successfully deployed at three other institutions with similar accuracy. DISCUSSION: A multi-modal strategy incorporating optical character recognition and natural language processing may increase the number of cases identified while maintaining similar PPVs. Such algorithms, however, require that the needed information be embedded within clinical documents. CONCLUSION: We have demonstrated that algorithms to identify and characterize cataracts can be developed utilizing data collected via the EHR. These algorithms provide a high level of accuracy even when implemented across multiple EHRs and institutional boundaries.
Peggy L. Peissig, Luke V. Rasmussen, Richard L. Berg, James G. Linneman, Catherine A. McCarty, Carol Waudby, Joshua C. Denny, Russell A. Wilke, Jyotishman Pathak, David Carrell, Abel N. Kho, Justin Starren
J. Am. Medical Informatics Assoc.11
2010 An analytical approach to characterize morbidity profile dissimilarity between distinct cohorts using electronic medical records
Jonathan S. Schildcrout, Melissa A. Basford, Jill M. Pulley, Daniel R. Masys, Dan M. Roden, Deede Wang, Christopher G. Chute, Iftikhar J. Kullo, David Carrell, Peggy L. Peissig, Abel N. Kho, Joshua C. Denny
J. Biomed. Informatics9
2007 Research Paper: Patient Web Services Integrated with a Shared Medical Record: Patient Use and Satisfaction
abstract
OBJECTIVES: This study sought to describe the evolution, use, and user satisfaction of a patient Web site providing a shared medical record between patients and health professionals at Group Health Cooperative, a mixed-model health care financing and delivery organization based in Seattle, Washington. DESIGN: This study used a retrospective, serial, cross-sectional study from September 2002 through December 2005 and a mailed satisfaction survey of a random sampling of 2,002 patients. MEASUREMENTS: This study measured the adoption and use of a patient Web site (MyGroupHealth) from September 2002 through December 2005. RESULTS: As of December 2005, 25% (105,047) of all Group Health members had registered and completed an identification verification process enabling them to use all of the available services on MyGroupHealth. Identification verification was more common among patients receiving care in the Integrated Delivery System (33%) compared with patients receiving care in the network (7%). As of December 2005, unique monthly user rates per 1,000 adult members were the highest for review of medical test results (54 of 1,000), medication refills (44 of 1,000), after-visit-summaries (32 of 1,000), and patient-provider clinical messaging (31 of 1,000). The response rate for the patient satisfaction survey was 46% (n = 921); 94% of survey respondents were satisfied or very satisfied with MyGroupHealth overall. Patients reported highest satisfaction (satisfied or very satisfied) for medication refills (96%), patient-provider messaging (93%), and medical test results (86%). CONCLUSION: Use and satisfaction with MyGroupHealth were greatest for accessing services and information involving ongoing, active care and patient-provider communication. Tight integration of Web services with clinical information systems and patient-provider relationships may be important in meeting the needs of patients.
James D. Ralston, David Carrell, Robert Reid 0002, Melissa A. Anderson, Maureena Moran, James Hereford
J. Am. Medical Informatics Assoc.2
2006 Variation in Adoption Rates of a Patient Web Portal with a Shared Medical Record by Age, Gender, and Morbidity Level
David Carrell, James D. Ralston
AMIA1
2006 Use and Satisfaction of a Patient Web Portal with a Shared Medical Record between Patients and Providers
James D. Ralston, James Hereford, David Carrell, Maureena Moran
AMIA3
2005 Messages, Strands and Threads: Measuring Use of Electronic Patient-Provider Messaging
David Carrell, James D. Ralston
AMIA1