David A. Hanauer

dblp:97/9193 · DBLP profile ↗
← Back
39ranked-venue papers
13as first author
9since 2021 · last 2026
0000-0001-6931-3791ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 36 · 13 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Opportunities for informatics to improve patient experiences: observations and reflections of ACMI fellows
abstract
OBJECTIVES: We report on findings from a meeting convened by the American College of Medical Informatics (ACMI) to characterize aspects of the patient experience that could be improved using informatics. MATERIALS AND METHODS: The American College of Medical Informatics fellows were invited to share their experiences as patients and suggest informatics approaches that may improve the patient experience. RESULTS: We identified 4 themes: (1) getting the right care, (2) data sharing and data interoperability, (3) guiding low-cost evaluations, and (4) predictive analytics. DISCUSSION: Despite widespread adoption of health IT, patient experiences remain far from optimal. CONCLUSION: The American College of Medical Informatics fellows identified informatics approaches, applications, and research areas that have the potential to improve patient experiences with health care systems.
Howard R. Strasberg, Edward P. Hoffer, Ross Koppel, Kevin B. Johnson, William M. Tierney, Geoffrey W. Rutledge, Elmer V. Bernstam, Jos Aarts, Marion J. Ball, Douglas S. Bell, Bernd Blobel, Suzanne Boren, Iain E. Buchan, James J. Cimino, Lawrence M. Fagan, James Geller, María Adela Grando, David A. Hanauer, William R. Hogan, Andrew S. Kanter, Bonnie Kaplan, Casimir A. Kulikowski, Albert Lai, David McCallie, Vimla Patel, Wanda Pratt, Sarah Collins Rossetti, Edward H. Shortliffe, Hardeep Singh 0005, Dean F. Sittig, William W. Stead, Kim M. Unertl, Mark G. Weiner, Kai Zheng 0002
J. Am. Medical Informatics Assoc.18
2025 Impacts of sample weighting on transferability of risk prediction models across EHR-Linked biobanks with different recruitment strategies
Maxwell Salvatore, Alison M. Mondul, Christopher R. Friese, David A. Hanauer, Hua Xu 0001, Celeste Leigh Pearce, Bhramar Mukherjee
J. Biomed. Informatics4
2024 To weight or not to weight? The effect of selection bias in 3 large electronic health record-linked biobanks and recommendations for practice
abstract
OBJECTIVES: To develop recommendations regarding the use of weights to reduce selection bias for commonly performed analyses using electronic health record (EHR)-linked biobank data. MATERIALS AND METHODS: We mapped diagnosis (ICD code) data to standardized phecodes from 3 EHR-linked biobanks with varying recruitment strategies: All of Us (AOU; n = 244 071), Michigan Genomics Initiative (MGI; n = 81 243), and UK Biobank (UKB; n = 401 167). Using 2019 National Health Interview Survey data, we constructed selection weights for AOU and MGI to represent the US adult population more. We used weights previously developed for UKB to represent the UKB-eligible population. We conducted 4 common analyses comparing unweighted and weighted results. RESULTS: For AOU and MGI, estimated phecode prevalences decreased after weighting (weighted-unweighted median phecode prevalence ratio [MPR]: 0.82 and 0.61), while UKB estimates increased (MPR: 1.06). Weighting minimally impacted latent phenome dimensionality estimation. Comparing weighted versus unweighted phenome-wide association study for colorectal cancer, the strongest associations remained unaltered, with considerable overlap in significant hits. Weighting affected the estimated log-odds ratio for sex and colorectal cancer to align more closely with national registry-based estimates. DISCUSSION: Weighting had a limited impact on dimensionality estimation and large-scale hypothesis testing but impacted prevalence and association estimation. When interested in estimating effect size, specific signals from untargeted association analyses should be followed up by weighted analysis. CONCLUSION: EHR-linked biobanks should report recruitment and selection mechanisms and provide selection weights with defined target populations. Researchers should consider their intended estimands, specify source and target populations, and weight EHR-linked biobank analyses accordingly.
Maxwell Salvatore, Ritoban Kundu, Christopher R. Friese, Seunggeun Lee, Lars G. Fritsche, Alison M. Mondul, David A. Hanauer, Celeste Leigh Pearce, Bhramar Mukherjee
J. Am. Medical Informatics Assoc.8
2023 Informative missingness: What can we learn from patterns in missing laboratory data in the electronic health record?
Amelia L. M. Tan, Emily J. Getzen, Meghan Hutch, Zachary H. Strasser, Alba Gutiérrez-Sacristán, Trang T. Le, Arianna Dagliati, Michele Morris, David A. Hanauer, Bertrand Moal, Clara-Lea Bonzel, William Yuan, Lorenzo Chiudinelli, Priyam Das, Harrison G. Zhang, Bruce J. Aronow, Paul Avillach, Gabriel A. Brat, Tianxi Cai, Chuan Hong, William G. La Cava, He Hooi Will Loh, Yuan Luo 0001, Shawn N. Murphy, Kee Yuan Hgiam, Gilbert S. Omenn, Lav P. Patel, Malarkodi J. Samayamuthu, Emily R. Shriver, Zahra Shakeri Hossein Abad, Byorn W. L. Tan, Shyam Visweswaran, Griffin M. Weber, Zongqi Xia, Bertrand Verdy, Qi Long, Danielle L. Mowery, John H. Holmes
J. Biomed. Informatics9
2022 Synergies between centralized and federated approaches to data quality: a report from the national COVID cohort collaborative
abstract
OBJECTIVE: In response to COVID-19, the informatics community united to aggregate as much clinical data as possible to characterize this new disease and reduce its impact through collaborative analytics. The National COVID Cohort Collaborative (N3C) is now the largest publicly available HIPAA limited dataset in US history with over 6.4 million patients and is a testament to a partnership of over 100 organizations. MATERIALS AND METHODS: We developed a pipeline for ingesting, harmonizing, and centralizing data from 56 contributing data partners using 4 federated Common Data Models. N3C data quality (DQ) review involves both automated and manual procedures. In the process, several DQ heuristics were discovered in our centralized context, both within the pipeline and during downstream project-based analysis. Feedback to the sites led to many local and centralized DQ improvements. RESULTS: Beyond well-recognized DQ findings, we discovered 15 heuristics relating to source Common Data Model conformance, demographics, COVID tests, conditions, encounters, measurements, observations, coding completeness, and fitness for use. Of 56 sites, 37 sites (66%) demonstrated issues through these heuristics. These 37 sites demonstrated improvement after receiving feedback. DISCUSSION: We encountered site-to-site differences in DQ which would have been challenging to discover using federated checks alone. We have demonstrated that centralized DQ benchmarking reveals unique opportunities for DQ improvement that will support improved research analytics locally and in aggregate. CONCLUSION: By combining rapid, continual assessment of DQ with a large volume of multisite data, it is possible to support more nuanced scientific questions with the scale and rigor that they require.
Emily R. Pfaff, Andrew T. Girvin, Davera Gabriel, Kristin Kostka, Michele Morris, Matvey Palchuk, Harold P. Lehmann, Benjamin R. C. Amor, Mark Bissell, Katie R. Bradwell, Sigfried Gold, Stephanie S. Hong, Johanna Loomba, Amin Manna, Julie A. McMurry, Emily Niehaus, Nabeel Qureshi, Anita Walden, Xiaohan Tanner Zhang, Richard L. Zhu, Richard A. Moffitt, Christopher G. Chute, William G. Adams, Shaymaa Al-Shukri, Alfred Anzalone, Ahmad Baghal, Tellen D. Bennett, Elmer V. Bernstam, Mark M. Bissell, Brian Bush, Thomas R. Campion Jr., Victor Castro, Jack Chang, Deepa D. Chaudhari, Wenjin Chen, San Chu, James J. Cimino, Keith A. Crandall, Mark Crooks, Sara J. Deakyne Davies, John Dipalazzo, David A. Dorr, Daniel Eckrich, Sarah E. Eltinge, Daniel G. Fort, Georgiy Golovko, Snehil Gupta, Melissa A. Haendel, Janos G. Hajagos, David A. Hanauer, Brett M. Harnett, Ronald Horswell, Nancy Huang, Steven G. Johnson, Michael Kahn, Kamil Khanipov, Curtis Kieler, Katherine Ruiz De Luzuriaga, Sarah E. Maidlow, Ashley Martinez, Jomol Mathew, James C. McClay, Gabriel McMahan, Brian Melancon, Stéphane M. Meystre, Lucio Miele, Hiroki Morizono, Ray Pablo, Lav P. Patel, Jimmy Phuong, Daniel J. Popham, Claudia P. Pulgarin, Indra Neil Sarkar, Nancy Sazo, Soko Setoguchi, Selvin Soby, Sirisha Surampalli, Christine Suver, Uma Maheswara Reddy Vangala, Shyam Visweswaran, James von Oehsen, Kellie M. Walters, Laura K. Wiley, David A. Williams, Adrian H. Zai
J. Am. Medical Informatics Assoc.50
2022 SurvMaximin: Robust federated approach to transporting survival risk prediction models
Harrison G. Zhang, Xin Xiong 0006, Chuan Hong, Griffin M. Weber, Gabriel A. Brat, Clara-Lea Bonzel, Yuan Luo 0001, Rui Duan 0004, Nathan P. Palmer, Meghan Hutch, Alba Gutiérrez-Sacristán, Riccardo Bellazzi, Luca Chiovato, Kelly Cho, Arianna Dagliati, Hossein Estiri, Noelia García-Barrio, Romain Griffier, David A. Hanauer, Yuk-Lam Ho, John H. Holmes, Mark S. Keller, Jeffrey G. Klann, Sehi L'Yi, Sara Lozano-Zahonero, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Bertrand Moal, Jason H. Moore, Michele Morris, Danielle L. Mowery, Shawn N. Murphy, Antoine Neuraz, Kee Yuan Ngiam, Gilbert S. Omenn, Lav P. Patel, Miguel Pedrera-Jiménez, Andrea Prunotto, Malarkodi J. Samayamuthu, Fernando J. Sanz Vidorreta, Emily Schriver, Petra Schubert, Pablo Serrano-Balazote, Andrew M. South, Amelia L. M. Tan, Byorn W. L. Tan, Valentina Tibollo, Patric Tippmann, Shyam Visweswaran, Zongqi Xia, William Yuan, Daniela Zöller, Isaac S. Kohane, Paul Avillach, Zijian Guo 0003, Tianxi Cai
J. Biomed. Informatics20
2021 Developing a standardized protocol for computational sentiment analysis research using health-related social media data
abstract
OBJECTIVE: Sentiment analysis is a popular tool for analyzing health-related social media content. However, existing studies exhibit numerous methodological issues and inconsistencies with respect to research design and results reporting, which could lead to biased data, imprecise or incorrect conclusions, or incomparable results across studies. This article reports a systematic analysis of the literature with respect to such issues. The objective was to develop a standardized protocol for improving the research validity and comparability of results in future relevant studies. MATERIALS AND METHODS: We developed the Protocol of Analysis of senTiment in Health (PATH) based on a systematic review that analyzed common research design choices and how such choices were made, or reported, among eligible studies published 2010-2019. RESULTS: Of 409 articles screened, 89 met the inclusion criteria. A total of 16 distinctive research design choices were identified, 9 of which have significant methodological or reporting inconsistencies among the articles reviewed, ranging from how relevance of study data was determined to how the sentiment analysis tool selected was validated. Based on this result, we developed the PATH protocol that encompasses all these distinctive design choices and highlights the ones for which careful consideration and detailed reporting are particularly warranted. CONCLUSIONS: A substantial degree of methodological and reporting inconsistencies exist in the extant literature that applied sentiment analysis to analyzing health-related social media data. The PATH protocol developed through this research may contribute to mitigating such issues in future relevant studies.
Tingjue Yin, Zhaoxian Hu, Yunan Chen 0001, David A. Hanauer, Kai Zheng 0002
J. Am. Medical Informatics Assoc.5
2021 Validation of an internationally derived patient severity phenotype to support COVID-19 analytics from electronic health record data
abstract
OBJECTIVE: The Consortium for Clinical Characterization of COVID-19 by EHR (4CE) is an international collaboration addressing coronavirus disease 2019 (COVID-19) with federated analyses of electronic health record (EHR) data. We sought to develop and validate a computable phenotype for COVID-19 severity. MATERIALS AND METHODS: Twelve 4CE sites participated. First, we developed an EHR-based severity phenotype consisting of 6 code classes, and we validated it on patient hospitalization data from the 12 4CE clinical sites against the outcomes of intensive care unit (ICU) admission and/or death. We also piloted an alternative machine learning approach and compared selected predictors of severity with the 4CE phenotype at 1 site. RESULTS: The full 4CE severity phenotype had pooled sensitivity of 0.73 and specificity 0.83 for the combined outcome of ICU admission and/or death. The sensitivity of individual code categories for acuity had high variability-up to 0.65 across sites. At one pilot site, the expert-derived phenotype had mean area under the curve of 0.903 (95% confidence interval, 0.886-0.921), compared with an area under the curve of 0.956 (95% confidence interval, 0.952-0.959) for the machine learning approach. Billing codes were poor proxies of ICU admission, with as low as 49% precision and recall compared with chart review. DISCUSSION: We developed a severity phenotype using 6 code classes that proved resilient to coding variability across international institutions. In contrast, machine learning approaches may overfit hospital-specific orders. Manual chart review revealed discrepancies even in the gold-standard outcomes, possibly owing to heterogeneous pandemic conditions. CONCLUSIONS: We developed an EHR-based severity phenotype for COVID-19 in hospitalized patients and validated it at 12 international sites.
Jeffrey G. Klann, Hossein Estiri, Griffin M. Weber, Bertrand Moal, Paul Avillach, Chuan Hong, Amelia L. M. Tan, Brett K. Beaulieu-Jones, Victor M. Castro, Thomas Maulhardt, Alon Geva, Alberto Malovini, Andrew M. South, Shyam Visweswaran, Michele Morris, Malarkodi J. Samayamuthu, Gilbert S. Omenn, Kee Yuan Ngiam, Kenneth D. Mandl, Martin Boeker, Karen L. Olson, Danielle L. Mowery, Robert W. Follett, David A. Hanauer, Riccardo Bellazzi, Jason H. Moore, Ne-Hooi Will Loh, Douglas S. Bell, Kavishwar B. Wagholikar, Luca Chiovato, Valentina Tibollo, Siegbert Rieg, Anthony L. L. J. Li, Vianney Jouhet, Emily Schriver, Zongqi Xia, Meghan Hutch, Yuan Luo 0001, Isaac S. Kohane, Gabriel A. Brat, Shawn N. Murphy
J. Am. Medical Informatics Assoc.24
2021 Phenotype risk scores (PheRS) for pancreatic cancer using time-stamped electronic health record data: Discovery and validation in two large biobanks
abstract
BACKGROUND: Traditional methods for disease risk prediction and assessment, such as diagnostic tests using serum, urine, blood, saliva or imaging biomarkers, have been important for identifying high-risk individuals for many diseases, leading to early detection and improved survival. For pancreatic cancer, traditional methods for screening have been largely unsuccessful in identifying high-risk individuals in advance of disease progression leading to high mortality and poor survival. Electronic health records (EHR) linked to genetic profiles provide an opportunity to integrate multiple sources of patient information for risk prediction and stratification. We leverage a constellation of temporally associated diagnoses available in the EHR to construct a summary risk score, called a phenotype risk score (PheRS), for identifying individuals at high-risk for having pancreatic cancer. The proposed PheRS approach incorporates the time with respect to disease onset into the prediction framework. We combine and contrast the PheRS with more well-known measures of inherited susceptibility, namely, the polygenic risk scores (PRS) for prediction of pancreatic cancer. METHODOLOGY: We first calculated pairwise, unadjusted associations between pancreatic cancer diagnosis and all possible other diagnoses across the medical phenome. We call these pairwise associations co-occurrences. After accounting for cross-phenotype correlations, the multivariable association estimates from a subset of relatively independent diagnoses were used to create a weighted sum PheRS. We constructed time-restricted risk scores using data from 38,359 participants in the Michigan Genomics Initiative (MGI) based on the diagnoses contained in the EHR at 0, 1, 2, and 5 years prior to the target pancreatic cancer diagnosis. The PheRS was assessed for predictability in the UK Biobank (UKB). We tested the relative contribution of PheRS when added to a model containing a summary measure of inherited genetic susceptibility (PRS) plus other covariates like age, sex, smoking status, drinking status, and body mass index (BMI). RESULTS: Our exploration of co-occurrence patterns identified expected associations while also revealing unexpected relationships that may warrant closer attention. Solely using the pancreatic cancer PheRS at 5 years before the target diagnoses yielded an AUC of 0.60 (95% CI = [0.58, 0.62]) in UKB. A larger predictive model including PheRS, PRS, and the covariates at the 5-year threshold achieved an AUC of 0.74 (95% CI = [0.72, 0.76]) in UKB. We note that PheRS does contribute independently in the joint model. Finally, scores at the top percentiles of the PheRS distribution demonstrated promise in terms of risk stratification. Scores in the top 2% were 10.20 (95% CI = [9.34, 12.99]) times more likely to identify cases than those in the bottom 98% in UKB at the 5-year threshold prior to pancreatic cancer diagnosis. CONCLUSIONS: We developed a framework for creating a time-restricted PheRS from EHR data for pancreatic cancer using the rich information content of a medical phenome. In addition to identifying hypothesis-generating associations for future research, this PheRS demonstrates a potentially important contribution in identifying high-risk individuals, even after adjusting for PRS for pancreatic cancer and other traditional epidemiologic covariates. The methods are generalizable to other phenotypic traits.
Maxwell Salvatore, Lauren J. Beesley, Lars G. Fritsche, David A. Hanauer, Alison M. Mondul, Celeste Leigh Pearce, Bhramar Mukherjee
J. Biomed. Informatics4
2018 Comprehensive process model of clinical information interaction in primary care: results of a "best-fit" framework synthesis
abstract
Objective: To describe a new, comprehensive process model of clinical information interaction in primary care (Clinical Information Interaction Model, or CIIM) based on a systematic synthesis of published research. Materials and Methods: We used the "best fit" framework synthesis approach. Searches were performed in PubMed, Embase, the Cumulative Index to Nursing and Allied Health Literature (CINAHL), PsycINFO, Library and Information Science Abstracts, Library, Information Science and Technology Abstracts, and Engineering Village. Two authors reviewed articles according to inclusion and exclusion criteria. Data abstraction and content analysis of 443 published papers were used to create a model in which every element was supported by empirical research. Results: The CIIM documents how primary care clinicians interact with information as they make point-of-care clinical decisions. The model highlights 3 major process components: (1) context, (2) activity (usual and contingent), and (3) influence. Usual activities include information processing, source-user interaction, information evaluation, selection of information, information use, clinical reasoning, and clinical decisions. Clinician characteristics, patient behaviors, and other professionals influence the process. Discussion: The CIIM depicts the complete process of information interaction, enabling a grasp of relationships previously difficult to discern. The CIIM suggests potentially helpful functionality for clinical decision support systems (CDSSs) to support primary care, including a greater focus on information processing and use. The CIIM also documents the role of influence in clinical information interaction; influencers may affect the success of CDSS implementations. Conclusion: The CIIM offers a new framework for achieving CDSS workflow integration and new directions for CDSS design that can support the work of diverse primary care clinicians.
Tiffany C. Veinot, Charles R. Senteio, David A. Hanauer, Julie C. Lowery
J. Am. Medical Informatics Assoc.3
2017 Two-year longitudinal assessment of physicians' perceptions after replacement of a longstanding homegrown electronic health record: does a J-curve of satisfaction really exist?
abstract
This report describes a 2-year prospective, longitudinal survey of attending physicians in 3 clinical areas (family medicine, general pediatrics, internal medicine) who experienced a transition from a homegrown electronic health record (EHR) to a vendor EHR. Participants were already highly familiar with using EHRs. Data were collected 1 month before and 3, 6, 13, and 25 months post implementation. Our primary goal was to determine if perceptions followed a J-curve pattern in which they initially dropped but eventually surpassed baseline measures. A J-curve was not found for any measures, including workflow, safety, communication, and satisfaction. Only the reminders and alerts measure dropped and then returned to baseline (U-curve); a few remained flatlined. Most dropped and remained below baseline (L-curve). The only measure that remained above baseline was documenting in the exam room with the patient. This study adds to the literature about current controversies surrounding EHR adoption and physician satisfaction.
David A. Hanauer, Greta L. Branford, Grant Greenberg, Sharon Kileny, Mick P. Couper, Kai Zheng 0002, Sung Won Choi
J. Am. Medical Informatics Assoc.1
2017 Development and empirical user-centered evaluation of semantically-based query recommendation for an electronic health record search engine
David A. Hanauer, Danny T. Y. Wu, Qiaozhu Mei, Katherine B. Murkowski-Steffy, V. G. Vinod Vydiswaran, Kai Zheng 0002
J. Biomed. Informatics1
2016 Identifying unmet informational needs in the inpatient setting to increase patient and caregiver engagement in the context of pediatric hematopoietic stem cell transplantation
abstract
BACKGROUND: Patient-centered care has been shown to improve patient outcomes, satisfaction, and engagement. However, there is a paucity of research on patient-centered care in the inpatient setting, including an understanding of unmet informational needs that may be limiting patient engagement. Pediatric hematopoietic stem cell transplantation (HSCT) represents an ideal patient population for elucidating unmet informational needs, due to the procedure's complexity and its requirement for caregiver involvement. METHODS: We conducted field observations and semi-structured interviews of pediatric HSCT caregivers and patients to identify informational challenges in the inpatient hospital setting. Data were analyzed using a thematic grounded theory approach. RESULTS: Three stages of the caregiving experience that could potentially be supported by a health information technology system, with the goal of enhancing patient/caregiver engagement, were identified: (1) navigating the health system and learning to communicate effectively with the healthcare team, (2) managing daily challenges of caregiving, and (3) transitioning from inpatient care to long-term outpatient management. DISCUSSION: We provide four practical recommendations to meet the informational needs of pediatric HSCT patients and caregivers: (1) provide patients/caregivers with real-time access to electronic health record data, (2) provide information about the clinical trials in which the patient is enrolled, (3) provide information about the patient's care team, and (4) properly prepare patients and caregivers for hospital discharge. CONCLUSION: Pediatric HSCT caregivers and patients have multiple informational needs that could be met with a health information technology system that integrates data from several sources, including electronic health records. Meeting these needs could reduce patients' and caregivers' anxiety surrounding the care process; reduce information asymmetry between caregivers/patients and providers; empower patients/caregivers to participate in the care process; and, ultimately, increase patient/caregiver engagement in the care process.
Elizabeth Kaziunas, David A. Hanauer, Mark S. Ackerman, Sung Won Choi
J. Am. Medical Informatics Assoc.2
2016 Assessing the readability of ClinicalTrials.gov
abstract
OBJECTIVE: ClinicalTrials.gov serves critical functions of disseminating trial information to the public and helping the trials recruit participants. This study assessed the readability of trial descriptions at ClinicalTrials.gov using multiple quantitative measures. MATERIALS AND METHODS: The analysis included all 165,988 trials registered at ClinicalTrials.gov as of April 30, 2014. To obtain benchmarks, the authors also analyzed 2 other medical corpora: (1) all 955 Health Topics articles from MedlinePlus and (2) a random sample of 100,000 clinician notes retrieved from an electronic health records system intended for conveying internal communication among medical professionals. The authors characterized each of the corpora using 4 surface metrics, and then applied 5 different scoring algorithms to assess their readability. The authors hypothesized that clinician notes would be most difficult to read, followed by trial descriptions and MedlinePlus Health Topics articles. RESULTS: Trial descriptions have the longest average sentence length (26.1 words) across all corpora; 65% of their words used are not covered by a basic medical English dictionary. In comparison, average sentence length of MedlinePlus Health Topics articles is 61% shorter, vocabulary size is 95% smaller, and dictionary coverage is 46% higher. All 5 scoring algorithms consistently rated CliniclTrials.gov trial descriptions the most difficult corpus to read, even harder than clinician notes. On average, it requires 18 years of education to properly understand these trial descriptions according to the results generated by the readability assessment algorithms. DISCUSSION AND CONCLUSION: Trial descriptions at CliniclTrials.gov are extremely difficult to read. Significant work is warranted to improve their readability in order to achieve CliniclTrials.gov's goal of facilitating information dissemination and subject recruitment.
Danny T. Y. Wu, David A. Hanauer, Qiaozhu Mei, Patricia M. Clark, Lawrence C. An, Joshua Proulx, Qing T. Zeng, V. G. Vinod Vydiswaran, Kevyn Collins-Thompson, Kai Zheng 0002
J. Am. Medical Informatics Assoc.2
2016 DREAM: Classification scheme for dialog acts in clinical research query mediation
Julia Hoxha, Praveen Chandar Ravichandran, Zhe He 0001, James J. Cimino, David A. Hanauer, Chunhua Weng
J. Biomed. Informatics5
2015 What Are Frequent Data Requests from Researchers? A Conceptual Model of Researchers' EHR Data Needs for Comparative Effectiveness Research
Gregory William Hruby, Praveen Chandar Ravichandran, Julia Hoxha, Eneida A. Mendonça, David A. Hanauer, Chunhua Weng
AMIA5
2015 Caveats of Using Social Media Data for Medical Research: A Report from a Study on Eye-Related Symptoms in Tweets
Yang Liu 0019, Tricia O'Brien, Esha Sondhi, Qiaozhu Mei, David A. Hanauer, Kai Zheng 0002
AMIA5
2015 Identifying Patterns Indicative of Copying/Pasting Behavior in Patient Generated Online Content
Tera L. Reynolds, V. G. Vinod Vydiswaran, Qiaozhu Mei, David A. Hanauer, Kai Zheng 0002
AMIA5
2015 Transition and Reflection in the Use of Health Information: The Case of Pediatric Bone Marrow Transplant Caregivers
abstract
The impact of health information on caregivers is of increasing interest to HCI/CSCW in designing systems to support the social and emotional dimensions of managing health. Drawing on an interview study, as well as corroborating data including a multi-year ethnography, we detail the practices of caregivers (particularly parents) in a bone marrow transplant (BMT) center. We examine the interconnections between information and emotion work performed by caregivers through a liminal lens, highlighting the BMT experience as a time of transition and reflection in which caregivers must quickly adapt to the new social world of the hospital and learn to manage a wide range of patient needs. The transition from parent to 'caregiver' is challenging, placing additional emotional burdens on the intensive information work for managing BMT. As a time of reflection, the BMT experience also provides an occasion for generative thinking and alternative approaches to health management. Our study findings call for health systems that reflect a design paradigm focused on 'transforming lives' rather than 'transferring information.'
Elizabeth Kaziunas, Ayse G. Büyüktür, Jasmine Jones, Sung Won Choi, David A. Hanauer, Mark S. Ackerman
CSCW5
2015 Paper versus EHR: simplistic comparisons may not capture current reality
abstract
The recent study by Taft and colleagues, which explores communication differences in paper versus electronic health records (EHRs), was both interesting and timely.1 EHRs are becoming a focal point for healthcare delivery in the US, yet the impact of EHRs on the patient-provider relationship remains poorly understood. Communication is at the heart of this relationship, and providers are concerned about the potential for EHRs to reduce the quality of their communications with patients.2,3 We would like to provide additional thoughts on Taft et al.’s reported findings and put them into a broader context. First, it is interesting to note that the authors found that EHRs fostered better communications with patients across nearly all measures. However, we wonder what might explain why a physician would greet a patient more warmly when walking into an exam room with a laptop computer vs. a paper chart. This suggests that the effect of EHRs vs. paper records must either be very strong and rapid or that there may be other factors at play. With regards to the overall premise of comparing paper to electronic charts – EHRs are capable of much more than paper records, but that capability is partly the reason why clinicians may perceive that patient communications have been impaired by EHRs. Clinicians in the exam room are taking on more tasks and interacting with the EHR in ways that were not possible with paper records. For example, tasks associated with an ambulatory EHR that do not have comparable actions in paper records include acknowledgment of medical assistant- or nurse-entered data via button clicks; e-prescribing, which enforces more conformity than paper prescriptions and may display numerous alerts that require review and confirmation; coding the encounter for billing purposes, which might previously have been handled by clerical staff; responding to reminders about immunizations and overdue tests (and subsequent order entry); documenting the encounter using dropdown menus, checkboxes, free-text entry, and many other modalities; and other documentation requirements that have resulted from new healthcare regulations, such as meaningful use. The Methods section of the Taft et al. article stated that their mock EHR was “styled after the Department of Veterans’ Affairs’ computerized patient record system,” but it is unclear how many of the aforementioned tasks were handled by the residents using the system in the study or even if the study’s mock EHR supported these complex functions. Practicing clinicians will readily recognize the challenges of handling these tasks, especially entering data through structured data entry forms and responding to computer-generated alerts, while interacting with patients.4 These additional tasks are not necessarily bad, and some only exist because EHRs enable them and because they are considered beneficial (eg, drug safety alerts), but they do present different challenges than paper records. Clinicians are certainly not required to perform all of the tasks required by the EHR while the patient is in the exam room, but time pressures often encourage them to do so. In summary, if residents in the Taft et al. experiment did not complete many of these additional tasks while interacting with patients, then the study scenarios may not represent realistic and typical EHR use. Further, while the authors clearly took great care to reduce potential biases, including not informing the residents or patient actors about the purpose of the study, it is not clear if the raters were also blinded to the study’s objectives. It would be interesting to see if the measures of residents’ communication skills would change if the raters were only provided with audio recordings of the physician-patient interactions, without the accompanying video footage. Furthermore, in this context, patients’ perceptions may be more relevant. While patient actors in the study were given a copy of the communication tool “so they could provide feedback to the residents’ supervisors at the end of the study,” their impressions about the residents’ communication skills were not reported. We look forward to future work on how EHR use impacts physician-patient communications, including comparisons of experienced clinicians using different EHRs in real practice settings. Some studies have noted communication differences between providers who engage in extensive in-room EHR use and those who use the EHR minimally during patient encounters.5 There is no standard etiquette about what does or does not constitute appropriate EHR use during an ambulatory patient encounter, but further study of EHR use and doctor-patient communication could help inform better EHR usage guidelines.
David A. Hanauer, Kai Zheng 0002
J. Am. Medical Informatics Assoc.1
2015 Supporting information retrieval from electronic health records: A report of University of Michigan's nine-year experience in developing and using the Electronic Medical Record Search Engine (EMERSE)
abstract
OBJECTIVE: This paper describes the University of Michigan's nine-year experience in developing and using a full-text search engine designed to facilitate information retrieval (IR) from narrative documents stored in electronic health records (EHRs). The system, called the Electronic Medical Record Search Engine (EMERSE), functions similar to Google but is equipped with special functionalities for handling challenges unique to retrieving information from medical text. MATERIALS AND METHODS: Key features that distinguish EMERSE from general-purpose search engines are discussed, with an emphasis on functions crucial to (1) improving medical IR performance and (2) assuring search quality and results consistency regardless of users' medical background, stage of training, or level of technical expertise. RESULTS: Since its initial deployment, EMERSE has been enthusiastically embraced by clinicians, administrators, and clinical and translational researchers. To date, the system has been used in supporting more than 750 research projects yielding 80 peer-reviewed publications. In several evaluation studies, EMERSE demonstrated very high levels of sensitivity and specificity in addition to greatly improved chart review efficiency. DISCUSSION: Increased availability of electronic data in healthcare does not automatically warrant increased availability of information. The success of EMERSE at our institution illustrates that free-text EHR search engines can be a valuable tool to help practitioners and researchers retrieve information from EHRs more effectively and efficiently, enabling critical tasks such as patient case synthesis and research data abstraction. CONCLUSION: EMERSE, available free of charge for academic use, represents a state-of-the-art medical IR tool with proven effectiveness and user acceptance.
David A. Hanauer, Qiaozhu Mei, James Law, Ritu Khanna, Kai Zheng 0002
J. Biomed. Informatics1
2014 What Is Asked in Clinical Data Request Forms? A Multi-site Thematic Analysis of Forms Towards Better Data Access Support
David A. Hanauer, Gregory William Hruby, Daniel Fort, Luke V. Rasmussen, Eneida A. Mendonça, Chunhua Weng
AMIA1
2014 Mining Consumer Health Vocabulary from Community-Generated Text
V. G. Vinod Vydiswaran, Qiaozhu Mei, David A. Hanauer, Kai Zheng 0002
AMIA3
2014 User-Created Groups in Health Forums: What Makes Them Special?
V. G. Vinod Vydiswaran, Yang Liu 0019, Kai Zheng 0002, David A. Hanauer, Qiaozhu Mei
ICWSM4
2014 Patient-initiated electronic health record amendment requests
abstract
BACKGROUND AND OBJECTIVE: Providing patients access to their medical records offers many potential benefits including identification and correction of errors. The process by which patients ask for changes to be made to their records is called an 'amendment request'. Little is known about the nature of such amendment requests and whether they result in modifications to the chart. METHODS: We conducted a qualitative content analysis of all patient-initiated amendment requests that our institution received over a 7-year period. Recurring themes were identified along three analytic dimensions: (1) clinical/documentation area, (2) patient motivation for making the request, and (3) outcome of the request. RESULTS: The dataset consisted of 818 distinct requests submitted by 181 patients. The majority of these requests (n=636, 77.8%) were made to rectify incorrect information and 49.7% of all requests were ultimately approved. In 6.6% of the requests, patients wanted valid information removed from their record, 27.8% of which were approved. Among all of the patients requesting a copy of their chart, only a very small percentage (approximately 0.2%) submitted an amendment request. CONCLUSIONS: The low number of amendment requests may be due to inadequate awareness by patients about how to make changes to their records. To make this approach effective, it will be important to inform patients of their right to view and amend records and about the process for doing so. Increasing patient access to medical records could encourage patient participation in improving the accuracy of medical records; however, caution should be used.
David A. Hanauer, Rebecca Preib, Kai Zheng 0002, Sung Won Choi
J. Am. Medical Informatics Assoc.1
2014 Applying MetaMap to Medline for identifying novel associations in a large clinical dataset: a feasibility analysis
abstract
OBJECTIVE: We describe experiments designed to determine the feasibility of distinguishing known from novel associations based on a clinical dataset comprised of International Classification of Disease, V.9 (ICD-9) codes from 1.6 million patients by comparing them to associations of ICD-9 codes derived from 20.5 million Medline citations processed using MetaMap. Associations appearing only in the clinical dataset, but not in Medline citations, are potentially novel. METHODS: Pairwise associations of ICD-9 codes were independently identified in both the clinical and Medline datasets, which were then compared to quantify their degree of overlap. We also performed a manual review of a subset of the associations to validate how well MetaMap performed in identifying diagnoses mentioned in Medline citations that formed the basis of the Medline associations. RESULTS: The overlap of associations based on ICD-9 codes in the clinical and Medline datasets was low: only 6.6% of the 3.1 million associations found in the clinical dataset were also present in the Medline dataset. Further, a manual review of a subset of the associations that appeared in both datasets revealed that co-occurring diagnoses from Medline citations do not always represent clinically meaningful associations. DISCUSSION: Identifying novel associations derived from large clinical datasets remains challenging. Medline as a sole data source for existing knowledge may not be adequate to filter out widely known associations. CONCLUSIONS: In this study, novel associations were not readily identified. Further improvements in accuracy and relevance for tools such as MetaMap are needed to realize their expected utility.
David A. Hanauer, Mohammed Saeed 0001, Kai Zheng 0002, Qiaozhu Mei, Kerby Shedden, Alan R. Aronson, Naren Ramakrishnan
J. Am. Medical Informatics Assoc.1
2013 Location Bias of Identifiers in Clinical Narratives
David A. Hanauer, Qiaozhu Mei, Bradley A. Malin, Kai Zheng 0002
AMIA1
2013 Modeling temporal relationships in large scale clinical associations
abstract
OBJECTIVE: We describe an approach for modeling temporal relationships in a large scale association analysis of electronic health record data. The addition of temporal information can inform hypothesis generation and help to explain the relationships. We applied this approach on a dataset containing 41.2 million time-stamped International Classification of Diseases, Ninth Revision (ICD-9) codes from 1.6 million patients. METHODS: We performed two independent analyses including a pairwise association analysis using a χ(2) test and a temporal analysis using a binomial test. Data were visualized using network diagrams and reviewed for clinical significance. RESULTS: We found nearly 400 000 highly associated pairs of ICD-9 codes with varying numbers of strong temporal associations ranging from ≥1 day to ≥10 years apart. Most of the findings were not considered clinically novel, although some, such as an association between Helicobacter pylori infection and diabetes, have recently been reported in the literature. The temporal analysis in our large cohort, however, revealed that diabetes usually preceded the diagnoses of H pylori, raising questions about possible cause and effect. DISCUSSION: Such analyses have significant limitations, some of which are due to known problems with ICD-9 codes and others to potentially incomplete data even at a health system level. Nevertheless, large scale association analyses with temporal modeling can help provide a mechanism for novel discovery in support of hypothesis generation. CONCLUSIONS: Temporal relationships can provide an additional layer of meaning in identifying and interpreting clinical associations.
David A. Hanauer, Naren Ramakrishnan
J. Am. Medical Informatics Assoc.1
2012 Hedging their Mets: The Use of Uncertainty Terms in Clinical Documents and its Potential Implications when Sharing the Documents with Patients
David A. Hanauer, Yang Liu 0019, Qiaozhu Mei, Frank J. Manion, Ulysses J. Balis, Kai Zheng 0002
AMIA1
2012 tranSMART Supports a Post-GWAS Data Coordinating Center
Dan Stuart, John Harju, David A. Hanauer, Dina Aronzon, Raveen Sharma, Frank J. Manion, Haiping Xia, Carolyn Hutter, Stephen Gruber
AMIA4
2012 Cooperative documentation: the patient problem list as a nexus in electronic health records
abstract
The patient Problem List (PL) is a mandated documentation component of electronic health records supporting the longitudinal summarization of patient information in addition to facilitating the coordination of care by multidisciplinary medical teams. In this paper, we report an ethnographic study that examined the institutionalization of the PL. Specifically, we explored: (1) how different groups (primary care providers, inpatient hospitalists, specialists, and emergency doctors) perceived the purposes of the PL differently; (2) how these deviated perceptions might affect their use of the PL; and (3) how the technical design of the PL facilitated or hindered the clinical practices of these groups. We found significant ambiguity regarding the definition, benefits, and use of the PL across different groups. We also found that certain groups (e.g. primary care providers) had developed effective cooperative strategies regarding the use of the PL; however, suboptimal usage was common among other user types, which could have a profound impact on quality of care and safety. Based on these findings, we provide suggestions to improve the design of the PL, particularly on strengthening its support on longitudinal and cooperative clinical practices.
Xiaomu Zhou, Kai Zheng 0002, Mark S. Ackerman, David A. Hanauer
CSCW4
2011 Experiences with mining temporal event sequences from electronic medical records: initial successes and some challenges
abstract
The standardization and wider use of electronic medical records (EMR) creates opportunities for better understanding patterns of illness and care within and across medical systems. Our interest is in the temporal history of event codes embedded in patients' records, specifically investigating frequently occurring sequences of event codes across patients. In studying data from more than 1.6 million patient histories at the University of Michigan Health system we quickly realized that frequent sequences, while providing one level of data reduction, still constitute a serious analytical challenge as many involve alternate serializations of the same sets of codes. To further analyze these sequences, we designed an approach where a partial order is mined from frequent sequences of codes. We demonstrate an EMR mining system called EMRView that enables exploration of the precedence relationships to quickly identify and visualize partial order information encoded in key classes of patients. We demonstrate some important nuggets learned through our approach and also outline key challenges for future research based on our experiences.
Debprakash Patnaik, Patrick Butler, Naren Ramakrishnan, Laxmi Parida, Benjamin J. Keller, David A. Hanauer
KDD6
2011 Using the time and motion method to study clinical work processes and workflow: methodological inconsistencies and a call for standardized research
abstract
OBJECTIVE: To identify ways for improving the consistency of design, conduct, and results reporting of time and motion (T&M) research in health informatics. MATERIALS AND METHODS: We analyzed the commonalities and divergences of empirical studies published 1990-2010 that have applied the T&M approach to examine the impact of health IT implementation on clinical work processes and workflow. The analysis led to the development of a suggested 'checklist' intended to help future T&M research produce compatible and comparable results. We call this checklist STAMP (Suggested Time And Motion Procedures). RESULTS: STAMP outlines a minimum set of 29 data/ information elements organized into eight key areas, plus three supplemental elements contained in an 'Ancillary Data' area, that researchers may consider collecting and reporting in their future T&M endeavors. DISCUSSION: T&M is generally regarded as the most reliable approach for assessing the impact of health IT implementation on clinical work. However, there exist considerable inconsistencies in how previous T&M studies were conducted and/or how their results were reported, many of which do not seem necessary yet can have a significant impact on quality of research and generalisability of results. Therefore, we deem it is time to call for standards that can help improve the consistency of T&M research in health informatics. This study represents an initial attempt. CONCLUSION: We developed a suggested checklist to improve the methodological and results reporting consistency of T&M research, so that meaningful insights can be derived from across-study synthesis and health informatics, as a field, will be able to accumulate knowledge from these studies.
Kai Zheng 0002, Michael H. Guo, David A. Hanauer
J. Am. Medical Informatics Assoc.3
2011 Handling anticipated exceptions in clinical care: investigating clinician use of 'exit strategies' in an electronic health records system
abstract
Unpredictable yet frequently occurring exception situations pervade clinical care. Handling them properly often requires aberrant actions temporarily departing from normal practice. In this study, the authors investigated several exception-handling procedures provided in an electronic health records system for facilitating clinical documentation, which the authors refer to as 'data entry exit strategies.' Through a longitudinal analysis of computer-recorded usage data, the authors found that (1) utilization of the exit strategies was not affected by postimplementation system maturity or patient visit volume, suggesting clinicians' needs to 'exit' unwanted situations are persistent; and (2) clinician type and gender are strong predictors of exit-strategy usage. Drilldown analyses further revealed that the exit strategies were judiciously used and enabled actions that would be otherwise difficult or impossible. However, many data entries recorded via them could have been 'properly' documented, yet were not, and a considerable proportion containing temporary or incomplete information was never subsequently amended. These findings may have significant implications for the design of safer and more user-friendly point-of-care information systems for healthcare.
Kai Zheng 0002, David A. Hanauer, Rema Padman, Michael P. Johnson, Anwar A. Hussain, Wen Ye 0003, Xiaomu Zhou, Herbert S. Diamond
J. Am. Medical Informatics Assoc.2
2011 Collaborative search in electronic health records
abstract
OBJECTIVE: A full-text search engine can be a useful tool for augmenting the reuse value of unstructured narrative data stored in electronic health records (EHR). A prominent barrier to the effective utilization of such tools originates from users' lack of search expertise and/or medical-domain knowledge. To mitigate the issue, the authors experimented with a 'collaborative search' feature through a homegrown EHR search engine that allows users to preserve their search knowledge and share it with others. This feature was inspired by the success of many social information-foraging techniques used on the web that leverage users' collective wisdom to improve the quality and efficiency of information retrieval. DESIGN: The authors conducted an empirical evaluation study over a 4-year period. The user sample consisted of 451 academic researchers, medical practitioners, and hospital administrators. The data were analyzed using a social-network analysis to delineate the structure of the user collaboration networks that mediated the diffusion of knowledge of search. RESULTS: The users embraced the concept with considerable enthusiasm. About half of the EHR searches processed by the system (0.44 million) were based on stored search knowledge; 0.16 million utilized shared knowledge made available by other users. The social-network analysis results also suggest that the user-collaboration networks engendered by the collaborative search feature played an instrumental role in enabling the transfer of search knowledge across people and domains. CONCLUSION: Applying collaborative search, a social information-foraging technique popularly used on the web, may provide the potential to improve the quality and efficiency of information retrieval in healthcare.
Kai Zheng 0002, Qiaozhu Mei, David A. Hanauer
J. Am. Medical Informatics Assoc.3
2010 Quantifying the impact of health IT implementations on clinical workflow: a new methodological perspective
abstract
Health IT implementations often introduce radical changes to clinical work processes and workflow. Prior research investigating this effect has shown conflicting results. Recent time and motion studies have consistently found that this impact is negligible; whereas qualitative studies have repeatedly revealed negative end-user perceptions suggesting decreased efficiency and disrupted workflow. We speculate that this discrepancy may be due in part to the design of the time and motion studies, which is focused on measuring clinicians' 'time expenditures' among different clinical activities rather than inspecting clinical 'workflow' from the true 'flow of the work' perspective. In this paper, we present a set of new analytical methods consisting of workflow fragmentation assessments, pattern recognition, and data visualization, which are accordingly designed to uncover hidden regularities embedded in the flow of the work. Through an empirical study, we demonstrate the potential value of these new methods in enriching workflow analysis in clinical settings.
Kai Zheng 0002, Hilary M. Haftel, Ronald B. Hirschl, Michael O'Reilly, David A. Hanauer
J. Am. Medical Informatics Assoc.5
2006 EMERSE: The Electronic Medical Record Search Engine
David A. Hanauer
AMIA1
2006 EMERSE: The Electronic Medical Record Search Engine
David A. Hanauer
AMIA1
2003 Use of the Internet for Seeking Health Care Information among Young Adults
David A. Hanauer, Jennifer M. Fortin, Emily Dibble, Nananda F. Col
AMIA1