Enrico W. Coiera

dblp:62/3722 · DBLP profile ↗
← Back
79ranked-venue papers
20as first author
15since 2021 · last 2025
0000-0002-6444-6584ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 69 · 16 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 The real-world impact of artificial intelligence ethics frameworks across a decade in healthcare: a scoping review
abstract
OBJECTIVES: The number of ethical frameworks designed to guide artificial intelligence (AI) use has grown substantially over the past decade, yet their real-world effect remains unclear. We aimed to synthesize existing evidence to analyze the practical impact of AI ethics frameworks (AIEFs) operationalized in healthcare. MATERIALS AND METHODS: We conducted a scoping review across 4 academic databases (Ovid MEDLINE, Ovid Embase, Scopus, and Web of Science), Google, and Google Scholar from January 2014 to January 2025. Eligible studies reported primary research on the qualitative or quantitative impacts of AIEFs implemented in healthcare. Data synthesis was conducted via narrative review. RESULTS: Of 1807 records identified, 16 studies met inclusion criteria. These comprised 5 preliminary initiatives testing guidelines in practice, 5 case studies, 5 implementation studies, and a comparative case study. AIEFs were implemented: (1) to develop new AI governance structures and guidelines, (2) as ethical review assessment systems for adopting clinical AI technologies, and (3) as ethical "audit" tools for identifying ethical risks. Impact was reported through qualitative improvements to process measures such as improved trust in AI. No studies demonstrated a direct link between AIEFs and health-related outcome measures such as patient safety. DISCUSSION: AIEFs led to changes in organizational or clinical processes, including increased compliance with ethical standards. When embedded in governance, AIEFs improved oversight and evaluation, but audits were constrained by their reliance on organizational cooperation. CONCLUSION: Despite the proliferation of AIEFs over the past decade, their implementation in healthcare remains limited and impact on health outcomes unmeasured or underreported.
Anastasia Chan, Hania Rahimi-Ardabili, Wendy A. Rogers, Enrico W. Coiera
J. Am. Medical Informatics Assoc.4
2025 A model for intelligible interaction between agents that predict and explain
A. Baskar 0001, Ashwin Srinivasan 0001, Michael Bain 0001, Enrico W. Coiera
Mach. Learn.4
2023 The standard problem
abstract
OBJECTIVE: This article proposes a framework to support the scientific research of standards so that they can be better measured, evaluated, and designed. METHODS: Beginning with the notion of common models, the framework describes the general standard problem-the seeming impossibility of creating a singular, persistent, and definitive standard which is not subject to change over time in an open system. RESULTS: The standard problem arises from uncertainty driven by variations in operating context, standard quality, differences in implementation, and drift over time. As a result, fitting work using conformance services is needed to repair these gaps between a standard and what is required for real-world use. To guide standards design and repair, a framework for measuring performance in context is suggested, based on signal detection theory and technomarkers. Based on the type of common model in operation, different conformance strategies are identified: (1) Universal conformance (all agents access the same standard); (2) Mediated conformance (an interoperability layer supports heterogeneous agents); and (3) Localized conformance (autonomous adaptive agents manage their own needs). Conformance methods include incremental design, modular design, adaptors, and creating interactive and adaptive agents. DISCUSSION: Machine learning should have a major role in adaptive fitting. Research to guide the choice and design of conformance services may focus on the stability and homogeneity of shared tasks, and whether common models are shared ahead of time or adjusted at task time. CONCLUSION: This analysis conceptually decouples interoperability and standardization. While standards facilitate interoperability, interoperability is achievable without standardization.
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
2023 Post-implementation optimization of medication alerts in hospital computerized provider order entry systems: a scoping review
abstract
OBJECTIVES: A scoping review identified interventions for optimizing hospital medication alerts post-implementation, and characterized the methods used, the populations studied, and any effects of optimization. MATERIALS AND METHODS: A structured search was undertaken in the MEDLINE and Embase databases, from inception to August 2023. Articles providing sufficient information to determine whether an intervention was conducted to optimize alerts were included in the analysis. Snowball analysis was conducted to identify additional studies. RESULTS: Sixteen studies were identified. Most were based in the United States and used a wide range of clinical software. Many studies used inpatient cohorts and conducted more than one intervention during the trial period. Alert types studied included drug-drug interactions, drug dosage alerts, and drug allergy alerts. Six types of interventions were identified: alert inactivation, alert severity reclassification, information provision, use of contextual information, threshold adjustment, and encounter suppression. The majority of interventions decreased alert quantity and enhanced alert acceptance. Alert quantity decreased with alert inactivation by 1%-25.3%, and with alert severity reclassification by 1%-16.5% in 6 of 7 studies. Alert severity reclassification increased alert acceptance by 4.2%-50.2% and was associated with a 100% acceptance rate for high-severity alerts when implemented. Clinical errors reported in 4 studies were seen to remain stable or decrease. DISCUSSION: Post-implementation medication optimization interventions have positive effects for clinicians when applied in a variety of settings. Less well reported are the impacts of these interventions on the clinical care of patients, and how endpoints such as alert quantity contribute to changes in clinician and pharmacist perceptions of alert fatigue. CONCLUSION: Well conducted alert optimization can reduce alert fatigue by reducing overall alert quantity, improving clinical acceptance, and enhancing clinical utility.
Thomas Stephen Ledger, Kalissa Brooke-Cowden, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2023 More than algorithms: an analysis of safety events involving ML-enabled medical devices reported to the FDA
abstract
OBJECTIVE: To examine the real-world safety problems involving machine learning (ML)-enabled medical devices. MATERIALS AND METHODS: We analyzed 266 safety events involving approved ML medical devices reported to the US FDA's MAUDE program between 2015 and October 2021. Events were reviewed against an existing framework for safety problems with Health IT to identify whether a reported problem was due to the ML device (device problem) or its use, and key contributors to the problem. Consequences of events were also classified. RESULTS: Events described hazards with potential to harm (66%), actual harm (16%), consequences for healthcare delivery (9%), near misses that would have led to harm if not for intervention (4%), no harm or consequences (3%), and complaints (2%). While most events involved device problems (93%), use problems (7%) were 4 times more likely to harm (relative risk 4.2; 95% CI 2.5-7). Problems with data input to ML devices were the top contributor to events (82%). DISCUSSION: Much of what is known about ML safety comes from case studies and the theoretical limitations of ML. We contribute a systematic analysis of ML safety problems captured as part of the FDA's routine post-market surveillance. Most problems involved devices and concerned the acquisition of data for processing by algorithms. However, problems with the use of devices were more likely to harm. CONCLUSIONS: Safety problems with ML devices involve more than algorithms, highlighting the need for a whole-of-system approach to safe implementation with a special focus on how users interact with devices.
David Lyell, Ying Wang 0003, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.3
2023 Using automated methods to detect safety problems with health information technology: a scoping review
abstract
OBJECTIVE: To summarize the research literature evaluating automated methods for early detection of safety problems with health information technology (HIT). MATERIALS AND METHODS: We searched bibliographic databases including MEDLINE, ACM Digital, Embase, CINAHL Complete, PsycINFO, and Web of Science from January 2010 to June 2021 for studies evaluating the performance of automated methods to detect HIT problems. HIT problems were reviewed using an existing classification for safety concerns. Automated methods were categorized into rule-based, statistical, and machine learning methods, and their performance in detecting HIT problems was assessed. The review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta Analyses extension for Scoping Reviews statement. RESULTS: Of the 45 studies identified, the majority (n = 27, 60%) focused on detecting use errors involving electronic health records and order entry systems. Machine learning (n = 22) and statistical modeling (n = 17) were the most common methods. Unsupervised learning was used to detect use errors in laboratory test results, prescriptions, and patient records while supervised learning was used to detect technical errors arising from hardware or software issues. Statistical modeling was used to detect use errors, unauthorized access, and clinical decision support system malfunctions while rule-based methods primarily focused on use errors. CONCLUSIONS: A wide variety of rule-based, statistical, and machine learning methods have been applied to automate the detection of safety problems with HIT. Many opportunities remain to systematically study their application and effectiveness in real-world settings.
Didi Surian, Ying Wang 0003, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.3
2022 Weak label based Bayesian U-Net for optic disc segmentation in fundus images
Hao Xiong 0001, Sidong Liu, Roneel V. Sharan, Enrico W. Coiera, Shlomo Berkovsky
Artif. Intell. Medicine4
2022 Symbolic and Statistical Learning Approaches to Speech Summarization: A Scoping Review
Dana Rezazadegan, Shlomo Berkovsky, Juan C. Quiroz, Ahmet Baki Kocaballi, Ying Wang 0003, Liliana Laranjo, Enrico W. Coiera
Comput. Speech Lang.7
2022 What did you do to avoid the climate disaster? A call to arms for health informatics
abstract
The effects of human-induced climate change on our planet are already largely irreversible for many centuries1 and may remain so for at least 1000 years.2 If emissions continue to grow, their effects will trigger multiple critical tipping points and event cascades that will amplify climate effects in unpredicted ways.3 Just in 2022, we have seen flooding cover one-third of Pakistan, affecting 33 million people.4 India and Pakistan’s heatwave was the hottest yet on record.5 Recent years have seen historically extreme forest fires across North America, Europe, Asia, and Australia. Low-lying Pacific nations are slowly starting to disappear as sea levels rise, fed by melt waters from disappearing glaciers and sea ice, and coastal cities everywhere are at risk.6 The same emissions driving climate change are also affecting our health. Particulates in air pollution are likely responsible for around 300 000 lung cancer deaths globally,7 and the list of climate-induced health problems leading to poorer outcomes is depressingly long.8,9 Humanity is in trouble, and our way out is uncertain to say the least.
Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.1
2022 Family informatics
abstract
While families have a central role in shaping individual choices and behaviors, healthcare largely focuses on treating individuals or supporting self-care. However, a family is also a health unit. We argue that family informatics is a necessary evolution in scope of health informatics. To deal with the needs of individuals, we must ensure technologies account for the role of their families and may require new classes of digital service. Social networks can help conceptualize the structure, composition, and behavior of families. A family network can be seen as a multiagent system with distributed cognition. Digital tools can address family needs in (1) sensing and monitoring; (2) communicating and sharing; (3) deciding and acting; and (4) treating and preventing illness. Family informatics is inherently multidisciplinary and has the potential to address unresolved chronic health challenges such as obesity, mental health, and substance abuse, support acute health challenges, and to improve the capacity of individuals to manage their own health needs.
Enrico W. Coiera, Kathleen Yin, Roneel V. Sharan, Saba Akbar, Satya Vedantam, Hao Xiong 0001, Jenny Waldie, Annie Y. S. Lau
J. Am. Medical Informatics Assoc.1
2022 Digital health for climate change mitigation and response: a scoping review
abstract
OBJECTIVE: Climate change poses a major threat to the operation of global health systems, triggering large scale health events, and disrupting normal system operation. Digital health may have a role in the management of such challenges and in greenhouse gas emission reduction. This scoping review explores recent work on digital health responses and mitigation approaches to climate change. MATERIALS AND METHODS: We searched Medline up to February 11, 2022, using terms for digital health and climate change. Included articles were categorized into 3 application domains (mitigation, infectious disease, or environmental health risk management), and 6 technical tasks (data sensing, monitoring, electronic data capture, modeling, decision support, and communication). The review was PRISMA-ScR compliant. RESULTS: The 142 included publications reported a wide variety of research designs. Publication numbers have grown substantially in recent years, but few come from low- and middle-income countries. Digital health has the potential to reduce health system greenhouse gas emissions, for example by shifting to virtual services. It can assist in managing changing patterns of infectious diseases as well as environmental health events by timely detection, reducing exposure to risk factors, and facilitating the delivery of care to under-resourced areas. DISCUSSION: While digital health has real potential to help in managing climate change, research remains preliminary with little real-world evaluation. CONCLUSION: Significant acceleration in the quality and quantity of digital health climate change research is urgently needed, given the enormity of the global challenge.
Hania Rahimi-Ardabili, Farah Magrabi, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2022 Consumer workarounds during the COVID-19 pandemic: analysis and technology implications using the SAMR framework
abstract
OBJECTIVE: To understand the nature of health consumer self-management workarounds during the COVID-19 pandemic; to classify these workarounds using the Substitution, Augmentation, Modification, and Redefinition (SAMR) framework; and to see how digital tools had assisted these workarounds. MATERIALS AND METHODS: We assessed 15 self-managing elderly patients with Type 2 diabetes, multiple chronic comorbidities, and low digital literacy. Interviews were conducted during COVID-19 lockdowns in May-June 2020 and participants were asked about how their self-management had differed from before. Each instance of change in self-management were identified as consumer workarounds and were classified using the SAMR framework to assess the extent of change. We also identified instances where digital technology assisted with workarounds. RESULTS: Consumer workarounds in all SAMR levels were observed. Substitution, describing change in work quality or how basic information was communicated, was easy to make and involved digital tools that replaced face-to-face communications, such as the telephone. Augmentation, describing changes in task mechanisms that enhanced functional value, did not include any digital tools. Modification, which significantly altered task content and context, involved more complicated changes such as making video calls. Redefinition workarounds created tasks not previously required, such as using Google Home to remotely babysit grandchildren, had transformed daily routines. DISCUSSION AND CONCLUSION: Health consumer workarounds need further investigation as health consumers also use workarounds to bypass barriers during self-management. The SAMR framework had classified the health consumer workarounds during COVID, but the framework needs further refinement to include more aspects of workarounds.
Kathleen Yin, Enrico W. Coiera, Joshua Jung, Urvashi Rohilla, Annie Y. S. Lau
J. Am. Medical Informatics Assoc.2
2021 Replication studies in the clinical decision support literature-frequency, fidelity, and impact
abstract
OBJECTIVE: To assess the frequency, fidelity, and impact of replication studies in the clinical decision support system (CDSS) literature. MATERIALS AND METHODS: A PRISMA-compliant review identified CDSS replications across 28 health and biomedical informatics journals. Included articles were assessed for fidelity to the original study using 5 categories: Identical, Substitutable, In-class, Augmented, and Out-of-class; and 7 IMPISCO domains: Investigators (I), Method (M), Population (P), Intervention (I), Setting (S), Comparator (C), and Outcome (O). A fidelity score and heat map were generated using the ratings. RESULTS: From 4063 publications matching search criteria for CDSS research, only 12/4063 (0.3%) were ultimately identified as replications. Six articles replicated but could not reproduce the results of the Han et al (2005) CPOE study showing mortality increase and, over time, changed from truth testing to generalizing this result. Other replications successfully tested variants of CDSS technology (2/12) or validated measurement instruments (4/12). DISCUSSION: A replication rate of 3 in a thousand studies is low even by the low rates in other disciplines. Several new reporting methods were developed for this study, including the IMPISCO framework, fidelity scores, and fidelity heat maps. A reporting structure for clearly identifying replication research is also proposed. CONCLUSION: There is an urgent need to better characterize which core CDSS principles require replication, identify past replication data, and conduct missing replication studies. Attention to replication should improve the efficiency and effectiveness of CDSS research and avoiding potentially harmful trial and error technology deployment.
Enrico W. Coiera, Huong Ly Tong
J. Am. Medical Informatics Assoc.1
2021 In Memoriam. Safe, Sound and Profound: A Tribute to Prof. John Fox, PhD, FACMI, FIAHSI (1948-2021)
abstract
A Tribute to
María Adela Grando, Enrico W. Coiera, David Glasspool, Jeremy C. Wyatt, Mor Peleg
J. Biomed. Informatics2
2021 Prediction of anxiety disorders using a feature ensemble based bayesian neural network
Hao Xiong 0001, Shlomo Berkovsky, Mia Romano, Roneel V. Sharan, Sidong Liu, Enrico W. Coiera, Lauren F. McLellan
J. Biomed. Informatics6
2020 Safety concerns with consumer-facing mobile health applications and their consequences: a scoping review
abstract
OBJECTIVE: To summarize the research literature about safety concerns with consumer-facing health apps and their consequences. MATERIALS AND METHODS: We searched bibliographic databases including PubMed, Web of Science, Scopus, and Cochrane libraries from January 2013 to May 2019 for articles about health apps. Descriptive information about safety concerns and consequences were extracted and classified into natural categories. The review was conducted in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) statement. RESULTS: Of the 74 studies identified, the majority were reviews of a single or a group of similar apps (n = 66, 89%), nearly half related to disease management (n = 34, 46%). A total of 80 safety concerns were identified, 67 related to the quality of information presented including incorrect or incomplete information, variation in content, and incorrect or inappropriate response to consumer needs. The remaining 13 related to app functionality including gaps in features, lack of validation for user input, delayed processing, failure to respond to health dangers, and faulty alarms. Of the 52 reports of actual or potential consequences, 5 had potential for patient harm. We also identified 66 reports about gaps in app development, including the lack of expert involvement, poor evidence base, and poor validation. CONCLUSIONS: Safety of apps is an emerging public health issue. The available evidence shows that apps pose clinical risks to consumers. Involvement of consumers, regulators, and healthcare professionals in development and testing can improve quality. Additionally, mandatory reporting of safety concerns is needed to improve outcomes.
Saba Akbar, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.2
2020 Envisioning an artificial intelligence documentation assistant for future primary care consultations: A co-design study with general practitioners
abstract
OBJECTIVE: The study sought to understand the potential roles of a future artificial intelligence (AI) documentation assistant in primary care consultations and to identify implications for doctors, patients, healthcare system, and technology design from the perspective of general practitioners. MATERIALS AND METHODS: Co-design workshops with general practitioners were conducted. The workshops focused on (1) understanding the current consultation context and identifying existing problems, (2) ideating future solutions to these problems, and (3) discussing future roles for AI in primary care. The workshop activities included affinity diagramming, brainwriting, and video prototyping methods. The workshops were audio-recorded and transcribed verbatim. Inductive thematic analysis of the transcripts of conversations was performed. RESULTS: Two researchers facilitated 3 co-design workshops with 16 general practitioners. Three main themes emerged: professional autonomy, human-AI collaboration, and new models of care. Major implications identified within these themes included (1) concerns with medico-legal aspects arising from constant recording and accessibility of full consultation records, (2) future consultations taking place out of the exam rooms in a distributed system involving empowered patients, (3) human conversation and empathy remaining the core tasks of doctors in any future AI-enabled consultations, and (4) questioning the current focus of AI initiatives on improved efficiency as opposed to patient care. CONCLUSIONS: AI documentation assistants will likely to be integral to the future primary care consultations. However, these technologies will still need to be supervised by a human until strong evidence for reliable autonomous performance is available. Therefore, different human-AI collaboration models will need to be designed and evaluated to ensure patient safety, quality of care, doctor safety, and doctor autonomy.
Ahmet Baki Kocaballi, Kiran Ijaz, Liliana Laranjo, Juan C. Quiroz, Dana Rezazadegan, Huong Ly Tong, Simon Willcock, Shlomo Berkovsky, Enrico W. Coiera
J. Am. Medical Informatics Assoc.9
2020 Can Unified Medical Language System-based semantic representation improve automated identification of patient safety incident reports by type and severity?
abstract
OBJECTIVE: The study sought to evaluate the feasibility of using Unified Medical Language System (UMLS) semantic features for automated identification of reports about patient safety incidents by type and severity. MATERIALS AND METHODS: Binary support vector machine (SVM) classifier ensembles were trained and validated using balanced datasets of critical incident report texts (n_type = 2860, n_severity = 1160) collected from a state-wide reporting system. Generalizability was evaluated on different and independent hospital-level reporting system. Concepts were extracted from report narratives using the UMLS Metathesaurus, and their relevance and frequency were used as semantic features. Performance was evaluated by F-score, Hamming loss, and exact match score and was compared with SVM ensembles using bag-of-words (BOW) features on 3 testing datasets (type/severity: n_benchmark = 286/116, n_original = 444/4837, n_independent =6000/5950). RESULTS: SVMs using semantic features met or outperformed those based on BOW features to identify 10 different incident types (F-score [semantics/BOW]: benchmark = 82.6%/69.4%; original = 77.9%/68.8%; independent = 78.0%/67.4%) and extreme-risk events (F-score [semantics/BOW]: benchmark = 87.3%/87.3%; original = 25.5%/19.8%; independent = 49.6%/52.7%). For incident type, the exact match score for semantic classifiers was consistently higher than BOW across all test datasets (exact match [semantics/BOW]: benchmark = 48.9%/39.9%; original = 57.9%/44.4%; independent = 59.5%/34.9%). DISCUSSION: BOW representations are not ideal for the automated identification of incident reports because they do not account for text semantics. UMLS semantic representations are likely to better capture information in report narratives, and thus may explain their superior performance. CONCLUSIONS: UMLS-based semantic classifiers were effective in identifying incidents by type and extreme-risk events, providing better generalizability than classifiers using BOW.
Ying Wang 0003, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.2
2019 Understanding and Measuring User Experience in Conversational Interfaces
abstract
Abstract Although various methods have been developed to evaluate conversational interfaces, there has been a lack of methods specifically focusing on evaluating user experience. This paper reviews the understandings of user experience (UX) in conversational interfaces literature and examines the six questionnaires commonly used for evaluating conversational systems in order to assess the potential suitability of these questionnaires to measure different UX dimensions in that context. The method to examine the questionnaires involved developing an assessment framework for main UX dimensions with relevant attributes and coding the items in the questionnaires according to the framework. The results show that (i) the understandings of UX notably differed in literature; (ii) four questionnaires included assessment items, in varying extents, to measure hedonic, aesthetic and pragmatic dimensions of UX; (iii) while the dimension of affect was covered by two questionnaires, playfulness, motivation, and frustration dimensions were covered by one questionnaire only. The largest coverage of UX dimensions has been provided by the Subjective Assessment of Speech System Interfaces (SASSI). We recommend using multiple questionnaires to obtain a more complete measurement of user experience or improve the assessment of a particular UX dimension. RESEARCH HIGHLIGHTS Varying understandings of UX in conversational interfaces literature. A UX assessment framework with UX dimensions and their relevant attributes. Descriptions of the six main questionnaires for evaluating conversational interfaces. A comparison of the six questionnaires based on their coverage of UX dimensions.
Ahmet Baki Kocaballi, Liliana Laranjo, Enrico W. Coiera
Interact. Comput.3
2019 A network model of activities in primary care consultations
abstract
OBJECTIVE: The objective of this study is to characterize the dynamic structure of primary care consultations by identifying typical activities and their inter-relationships to inform the design of automated approaches to clinical documentation using natural language processing and summarization methods. MATERIALS AND METHODS: This is an observational study in Australian general practice involving 31 consultations with 4 primary care physicians. Consultations were audio-recorded, and computer interactions were recorded using screen capture. Physical interactions in consultation rooms were noted by observers. Brief interviews were conducted after consultations. Conversational transcripts were analyzed to identify different activities and their speech content as well as verbal cues signaling activity transitions. An activity transition analysis was then undertaken to generate a network of activities and transitions. RESULTS: Observed activity classes followed those described in well-known primary care consultation models. Activities were often fragmented across consultations, did not flow necessarily in a defined order, and the flow between activities was nonlinear. Modeling activities as a network revealed that discussing a patient's present complaint was the most central activity and was highly connected to medical history taking, physical examination, and assessment, forming a highly interrelated bundle. Family history, allergy, and investigation discussions were less connected suggesting less dependency on other activities. Clear verbal signs were often identifiable at transitions between activities. DISCUSSION: Primary care consultations do not appear to follow a classic linear model of defined information seeking activities; rather, they are fragmented, highly interdependent, and can be reactively triggered. CONCLUSION: The nonlinearity of activities has significant implications for the design of automated information capture. Whereas dictation systems generate literal translation of speech into text, speech-based clinical summary systems will need to link disparate information fragments, merge their content, and abstract coherent information summaries.
Ahmet Baki Kocaballi, Enrico W. Coiera, Huong Ly Tong, Sarah J. White, Juan C. Quiroz, Fahimeh Rezazadegan, Simon Willcock, Liliana Laranjo
J. Am. Medical Informatics Assoc.2
2019 Using convolutional neural networks to identify patient safety incident reports by type and severity
abstract
OBJECTIVE: To evaluate the feasibility of a convolutional neural network (CNN) with word embedding to identify the type and severity of patient safety incident reports. MATERIALS AND METHODS: A CNN with word embedding was applied to identify 10 incident types and 4 severity levels. Model training and validation used data sets (n_type = 2860, n_severity = 1160) collected from a statewide incident reporting system. Generalizability was evaluated using an independent hospital-level reporting system. CNN architectures were examined by varying layer size and hyperparameters. Performance was evaluated by F score, precision, recall, and compared to binary support vector machine (SVM) ensembles on 3 testing data sets (type/severity: n_benchmark = 286/116, n_original = 444/4837, n_independent = 6000/5950). RESULTS: A CNN with 6 layers was the most effective architecture, outperforming SVMs with better generalizability to identify incidents by type and severity. The CNN achieved high F scores (> 85%) across all test data sets when identifying common incident types including falls, medications, pressure injury, and aggression. When identifying common severity levels (medium/low), CNN outperformed SVMs, improving F scores by 11.9%-45.1% across all 3 test data sets. DISCUSSION: Automated identification of incident reports using machine learning is challenging because of a lack of large labelled training data sets and the unbalanced distribution of incident classes. The standard classification strategy is to build multiple binary classifiers and pool their predictions. CNNs can extract hierarchical features and assist in addressing class imbalance, which may explain their success in identifying incident report types. CONCLUSION: A CNN with word embedding was effective in identifying incidents by type and severity, providing better generalizability than SVMs.
Ying Wang 0003, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.2
2019 Tracking a moving user in indoor environments using Bluetooth low energy beacons
Didi Surian, Vitaliy Kim, Ranjeeta Menon, Adam G. Dunn, Vitali Sintchenko, Enrico W. Coiera
J. Biomed. Informatics6
2018 Safety concerns with consumer-facing mobile health applications and their consequences
Saba Akbar, Jessica A. Chen, Liliana Laranjo, Annie Y. S. Lau, Enrico W. Coiera, Farah Magrabi
AMIA5
2018 Technological Characteristics of Conversational Agents Used for Health-Related Purposes - A Systematic Review
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera
AMIA11
2018 Does health informatics have a replication crisis?
abstract
Objective: Many research fields, including psychology and basic medical sciences, struggle with poor reproducibility of reported studies. Biomedical and health informatics is unlikely to be immune to these challenges. This paper explores replication in informatics and the unique challenges the discipline faces. Methods: Narrative review of recent literature on research replication challenges. Results: While there is growing interest in re-analysis of existing data, experimental replication studies appear uncommon in informatics. Context effects are a particular challenge as they make ensuring replication fidelity difficult, and the same intervention will never quite reproduce the same result in different settings. Replication studies take many forms, trading-off testing validity of past findings against testing generalizability. Exact and partial replication designs emphasize testing validity while quasi and conceptual studies test generalizability of an underlying model or hypothesis with different methods or in a different setting. Conclusions: The cost of poor replication is a weakening in the quality of published research and the evidence-based foundation of health informatics. The benefits of replication include increased rigor in research, and the development of evaluation methods that distinguish the impact of context and the nonreproducibility of research. Taking replication seriously is essential if biomedical and health informatics is to be an evidence-based discipline.
Enrico W. Coiera, Elske Ammenwerth, Andrew Georgiou, Farah Magrabi
J. Am. Medical Informatics Assoc.1
2018 Conversational agents in healthcare: a systematic review
abstract
Objective: Our objective was to review the characteristics, current applications, and evaluation measures of conversational agents with unconstrained natural language input capabilities used for health-related purposes. Methods: We searched PubMed, Embase, CINAHL, PsycInfo, and ACM Digital using a predefined search strategy. Studies were included if they focused on consumers or healthcare professionals; involved a conversational agent using any unconstrained natural language input; and reported evaluation measures resulting from user interaction with the system. Studies were screened by independent reviewers and Cohen's kappa measured inter-coder agreement. Results: The database search retrieved 1513 citations; 17 articles (14 different conversational agents) met the inclusion criteria. Dialogue management strategies were mostly finite-state and frame-based (6 and 7 conversational agents, respectively); agent-based strategies were present in one type of system. Two studies were randomized controlled trials (RCTs), 1 was cross-sectional, and the remaining were quasi-experimental. Half of the conversational agents supported consumers with health tasks such as self-care. The only RCT evaluating the efficacy of a conversational agent found a significant effect in reducing depression symptoms (effect size d = 0.44, p = .04). Patient safety was rarely evaluated in the included studies. Conclusions: The use of conversational agents with unconstrained natural language input capabilities for health-related purposes is an emerging field of research, where the few published studies were mainly quasi-experimental, and rarely evaluated efficacy or safety. Future studies would benefit from more robust experimental designs and standardized reporting. Protocol Registration: The protocol for this systematic review is registered at PROSPERO with the number CRD42017065917.
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera
J. Am. Medical Informatics Assoc.11
2018 A shared latent space matrix factorisation method for recommending new trial evidence for systematic review updates
Didi Surian, Adam G. Dunn, Liat Orenstein, Rabia Bashir, Enrico W. Coiera, Florence T. Bourgeois
J. Biomed. Informatics5
2017 Engineering technology resilience through informatics safety science
abstract
With every year that passes, our relationship to information technology becomes more complex, and our dependence deeper. Technology is our great ally, promising greater efficiency and productivity. It also promises greater safety for our patients. However, this relationship with technology can sometimes be a brittle one. We can quickly cross a safety gap from a comfortable place where everything works well, to one where the limits of technology introduce new risks. Whether it is through a computer network failure, applying a system software patch, or a user accidentally clicking on the wrong patient name, it is surprisingly easy to move from safe to unsafe. As the footprint of technology across our health services has grown, so to by extrapolation, has the associated risk of technology harms to patients.1 It is the potential abruptness of this transition to increased risk of harm, this lack of graceful degradation in performance, and the silence accompanying degradation, that remain unsolved challenges to the effective use of information technology in healthcare.
Enrico W. Coiera, Farah Magrabi, Jan L. Talmon
J. Am. Medical Informatics Assoc.1
2017 Efficiency and safety of speech recognition for documentation in the electronic health record
abstract
OBJECTIVE: To compare the efficiency and safety of using speech recognition (SR) assisted clinical documentation within an electronic health record (EHR) system with use of keyboard and mouse (KBM). METHODS: Thirty-five emergency department clinicians undertook randomly allocated clinical documentation tasks using KBM or SR on a commercial EHR system. Tasks were simple or complex, and with or without interruption. Outcome measures included task completion times and observed errors. Errors were classed by their potential for patient harm. Error causes were classified as due to IT system/system integration, user interaction, comprehension, or as typographical. User-related errors could be by either omission or commission. RESULTS: Mean task completion times were 18.11% slower overall when using SR compared to KBM (P = .001), 16.95% slower for simple tasks (P = .050), and 18.40% slower for complex tasks (P = .009). Increased errors were observed with use of SR (KBM 32, SR 138) for both simple (KBM 9, SR 75; P < 0.001) and complex (KBM 23, SR 63; P < 0.001) tasks. Interruptions did not significantly affect task completion times or error rates for either modality. DISCUSSION: For clinical documentation, SR was slower and increased the risk of documentation errors, including errors with the potential to cause clinical harm compared to KBM. Some of the observed increase in errors may be due to suboptimal SR to EHR integration and workflow. CONCLUSION: Use of SR to drive interactive clinical documentation in the EHR requires careful evaluation. Current generation implementations may require significant development before they are safe and effective. Improving system integration and workflow, as well as SR accuracy and user-focused error correction strategies, may improve SR performance.
Tobias Hodgson, Farah Magrabi, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2017 Problems with health information technology and their effects on care delivery and patient outcomes: a systematic review
abstract
Objective: To systematically review studies reporting problems with information technology (IT) in health care and their effects on care delivery and patient outcomes. Materials and methods: We searched bibliographic databases including Scopus, PubMed, and Science Citation Index Expanded from January 2004 to December 2015 for studies reporting problems with IT and their effects. A framework called the information value chain, which connects technology use to final outcome, was used to assess how IT problems affect user interaction, information receipt, decision-making, care processes, and patient outcomes. The review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement. Results: Of the 34 studies identified, the majority ( n = 14, 41%) were analyses of incidents reported from 6 countries. There were 7 descriptive studies, 9 ethnographic studies, and 4 case reports. The types of IT problems were similar to those described in earlier classifications of safety problems associated with health IT. The frequency, scale, and severity of IT problems were not adequately captured within these studies. Use errors and poor user interfaces interfered with the receipt of information and led to errors of commission when making decisions. Clinical errors involving medications were well characterized. Issues with system functionality, including poor user interfaces and fragmented displays, delayed care delivery. Issues with system access, system configuration, and software updates also delayed care. In 18 studies (53%), IT problems were linked to patient harm and death. Near-miss events were reported in 10 studies (29%). Discussion and conclusion: The research evidence describing problems with health IT remains largely qualitative, and many opportunities remain to systematically study and quantify risks and benefits with regard to patient safety. The information value chain, when used in conjunction with existing classifications for health IT safety problems, can enhance measurement and should facilitate identification of the most significant risks to patient safety.
Mi Ok Kim, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.2
2017 Automation bias and verification complexity: a systematic review
abstract
INTRODUCTION: While potentially reducing decision errors, decision support systems can introduce new types of errors. Automation bias (AB) happens when users become overreliant on decision support, which reduces vigilance in information seeking and processing. Most research originates from the human factors literature, where the prevailing view is that AB occurs only in multitasking environments. OBJECTIVES: This review seeks to compare the human factors and health care literature, focusing on the apparent association of AB with multitasking and task complexity. DATA SOURCES: EMBASE, Medline, Compendex, Inspec, IEEE Xplore, Scopus, Web of Science, PsycINFO, and Business Source Premiere from 1983 to 2015. STUDY SELECTION: Evaluation studies where task execution was assisted by automation and resulted in errors were included. Participants needed to be able to verify automation correctness and perform the task manually. METHODS: Tasks were identified and grouped. Task and automation type and presence of multitasking were noted. Each task was rated for its verification complexity. RESULTS: Of 890 papers identified, 40 met the inclusion criteria; 6 were in health care. Contrary to the prevailing human factors view, AB was found in single tasks, typically involving diagnosis rather than monitoring, and with high verification complexity. LIMITATIONS: The literature is fragmented, with large discrepancies in how AB is reported. Few studies reported the statistical significance of AB compared to a control condition. CONCLUSION: AB appears to be associated with the degree of cognitive load experienced in decision tasks, and appears to not be uniquely associated with multitasking. Strategies to minimize AB might focus on cognitive load reduction.
David Lyell, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
2016 Real-time prediction of mortality, readmission, and length of stay using electronic health record data
abstract
OBJECTIVE: To develop a predictive model for real-time predictions of length of stay, mortality, and readmission for hospitalized patients using electronic health records (EHRs). MATERIALS AND METHODS: A Bayesian Network model was built to estimate the probability of a hospitalized patient being "at home," in the hospital, or dead for each of the next 7 days. The network utilizes patient-specific administrative and laboratory data and is updated each time a new pathology test result becomes available. Electronic health records from 32 634 patients admitted to a Sydney metropolitan hospital via the emergency department from July 2008 through December 2011 were used. The model was tested on 2011 data and trained on the data of earlier years. RESULTS: The model achieved an average daily accuracy of 80% and area under the receiving operating characteristic curve (AUROC) of 0.82. The model's predictive ability was highest within 24 hours from prediction (AUROC = 0.83) and decreased slightly with time. Death was the most predictable outcome with a daily average accuracy of 93% and AUROC of 0.84. DISCUSSION: We developed the first non-disease-specific model that simultaneously predicts remaining days of hospitalization, death, and readmission as part of the same outcome. By providing a future daily probability for each outcome class, we enable the visualization of future patient trajectories. Among these, it is possible to identify trajectories indicating expected discharge, expected continuing hospitalization, expected death, and possible readmission. CONCLUSIONS: Bayesian Networks can model EHRs to provide real-time forecasts for patient outcomes, which provide richer information than traditional independent point predictions of length of stay, death, or readmission, and can thus better support decision making.
Xiongcai Cai, Óscar Pérez, Enrico W. Coiera, Fernando Martín-Sánchez, Richard O. Day, David Roffe, Blanca Gallego
J. Am. Medical Informatics Assoc.3
2016 Risks and benefits of speech recognition for clinical documentation: a systematic review
abstract
OBJECTIVE: To review literature assessing the impact of speech recognition (SR) on clinical documentation. METHODS: Studies published prior to December 2014 reporting clinical documentation using SR were identified by searching Scopus, Compendex and Inspect, PubMed, and Google Scholar. Outcome variables analyzed included dictation and editing time, document turnaround time (TAT), SR accuracy, error rates per document, and economic benefit. Twenty-three articles met inclusion criteria from a pool of 441. RESULTS: Most studies compared SR to dictation and transcription (DT) in radiology, and heterogeneity across studies was high. Document editing time increased using SR compared to DT in four of six studies (+1876.47% to -16.50%). Dictation time similarly increased in three of five studies (+91.60% to -25.00%). TAT consistently improved using SR compared to DT (16.41% to 82.34%); across all studies the improvement was 0.90% per year. SR accuracy was reported in ten studies (88.90% to 96.00%) and appears to improve 0.03% per year as the technology matured. Mean number of errors per report increased using SR (0.05 to 6.66) compared to DT (0.02 to 0.40). Economic benefits were poorly reported. CONCLUSIONS: SR is steadily maturing and offers some advantages for clinical documentation. However, evidence supporting the use of SR is weak, and further investigation is required to assess the impact of SR on documentation error types, rates, and clinical outcomes.
Tobias Hodgson, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
2016 Measuring the effects of computer downtime on hospital pathology processes
Ying Wang 0003, Enrico W. Coiera, Blanca Gallego, Óscar Pérez, Mei-Sing Ong, Guy Tsafnat, David Roffe, Graham Jones, Farah Magrabi
J. Biomed. Informatics2
2014 Communication spaces
abstract
BACKGROUND AND OBJECTIVE: Annotations to physical workspaces such as signs and notes are ubiquitous. When densely annotated, work areas become communication spaces. This study aims to characterize the types and purpose of such annotations. METHODS: A qualitative observational study was undertaken in two wards and the radiology department of a 440-bed metropolitan teaching hospital. Images were purposefully sampled; 39 were analyzed after excluding inferior images. RESULTS: Annotation functions included signaling identity, location, capability, status, availability, and operation. They encoded data, rules or procedural descriptions. Most aggregated into groups that either created a workflow by referencing each other, supported a common workflow without reference to each other, or were heterogeneous, referring to many workflows. Higher-level assemblies of such groupings were also observed. DISCUSSION: Annotations make visible the gap between work done and the capability of a space to support work. Annotations are repairs of an environment, improving fitness for purpose, fixing inadequacy in design, or meeting emergent needs. Annotations thus record the missing information needed to undertake tasks, typically added post-implemented. Measuring annotation levels post-implementation could help assess the fit of technology to task. Physical and digital spaces could meet broader user needs by formally supporting user customization, 'programming through annotation'. Augmented reality systems could also directly support annotation, addressing existing information gaps, and enhancing work with context sensitive annotation. CONCLUSIONS: Communication spaces offer a model of how work unfolds. Annotations make visible local adaptation that makes technology fit for purpose post-implementation and suggest an important role for annotatable information systems and digital augmentation of the physical environment.
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
2014 Gene-disease association with literature based enrichment
Guy Tsafnat, Dennis Jasch, Agam Misra, Miew Keen Choong, Frank P. Y. Lin, Enrico W. Coiera
J. Biomed. Informatics6
2013 Using statistical text classification to identify health information technology incidents
abstract
OBJECTIVE: To examine the feasibility of using statistical text classification to automatically identify health information technology (HIT) incidents in the USA Food and Drug Administration (FDA) Manufacturer and User Facility Device Experience (MAUDE) database. DESIGN: We used a subset of 570 272 incidents including 1534 HIT incidents reported to MAUDE between 1 January 2008 and 1 July 2010. Text classifiers using regularized logistic regression were evaluated with both 'balanced' (50% HIT) and 'stratified' (0.297% HIT) datasets for training, validation, and testing. Dataset preparation, feature extraction, feature selection, cross-validation, classification, performance evaluation, and error analysis were performed iteratively to further improve the classifiers. Feature-selection techniques such as removing short words and stop words, stemming, lemmatization, and principal component analysis were examined. MEASUREMENTS: κ statistic, F1 score, precision and recall. RESULTS: Classification performance was similar on both the stratified (0.954 F1 score) and balanced (0.995 F1 score) datasets. Stemming was the most effective technique, reducing the feature set size to 79% while maintaining comparable performance. Training with balanced datasets improved recall (0.989) but reduced precision (0.165). CONCLUSIONS: Statistical text classification appears to be a feasible method for identifying HIT reports within large databases of incidents. Automated identification should enable more HIT problems to be detected, analyzed, and addressed in a timely manner. Semi-supervised learning may be necessary when applying machine learning to big data analysis of patient safety incidents and requires further investigation.
Kevin E. K. Chai, Stephen Anthony, Enrico W. Coiera, Farah Magrabi
J. Am. Medical Informatics Assoc.3
2013 Syndromic surveillance for health information system failures: a feasibility study
abstract
OBJECTIVE: To explore the applicability of a syndromic surveillance method to the early detection of health information technology (HIT) system failures. METHODS: A syndromic surveillance system was developed to monitor a laboratory information system at a tertiary hospital. Four indices were monitored: (1) total laboratory records being created; (2) total records with missing results; (3) average serum potassium results; and (4) total duplicated tests on a patient. The goal was to detect HIT system failures causing: data loss at the record level; data loss at the field level; erroneous data; and unintended duplication of data. Time-series models of the indices were constructed, and statistical process control charts were used to detect unexpected behaviors. The ability of the models to detect HIT system failures was evaluated using simulated failures, each lasting for 24 h, with error rates ranging from 1% to 35%. RESULTS: In detecting data loss at the record level, the model achieved a sensitivity of 0.26 when the simulated error rate was 1%, while maintaining a specificity of 0.98. Detection performance improved with increasing error rates, achieving a perfect sensitivity when the error rate was 35%. In the detection of missing results, erroneous serum potassium results and unintended repetition of tests, perfect sensitivity was attained when the error rate was as small as 5%. Decreasing the error rate to 1% resulted in a drop in sensitivity to 0.65-0.85. CONCLUSIONS: Syndromic surveillance methods can potentially be applied to monitor HIT systems, to facilitate the early detection of failures.
Mei-Sing Ong, Farah Magrabi, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2012 National and cross-border safety initiatives for health information technology
Farah Magrabi, Dean F. Sittig, Maureen Baker, Jan L. Talmon, Enrico W. Coiera
AMIA5
2012 The dangerous decade
abstract
Over the next 10 years, more information and communication technology (ICT) will be deployed in the health system than in its entire previous history. Systems will be larger in scope, more complex, and move from regional to national and supranational scale. Yet we are at roughly the same place the aviation industry was in the 1950s with respect to system safety. Even if ICT harm rates do not increase, increased ICT use will increase the absolute number of ICT related harms. Factors that could diminish ICT harm include adoption of common standards, technology maturity, better system development, testing, implementation and end user training. Factors that will increase harm rates include complexity and heterogeneity of systems and their interfaces, rapid implementation and poor training of users. Mitigating these harms will not be easy, as organizational inertia is likely to generate a hysteresis-like lag, where the paths to increase and decrease harm are not identical.
Enrico W. Coiera, Jos Aarts, Casimir A. Kulikowski
J. Am. Medical Informatics Assoc.1
2012 Impact of a web-based personally controlled health management system on influenza vaccination and health services utilization rates: a randomized controlled trial
abstract
OBJECTIVE: To assess the impact of a web-based personally controlled health management system (PCHMS) on the uptake of seasonal influenza vaccine and primary care service utilization among university students and staff. MATERIALS AND METHODS: A PCHMS called Healthy.me was developed and evaluated in a 2010 CONSORT-compliant two-group (6-month waitlist vs PCHMS) parallel randomized controlled trial (RCT) (allocation ratio 1:1). The PCHMS integrated an untethered personal health record with consumer care pathways, social forums, and messaging links with a health service provider. RESULTS: 742 university students and staff met inclusion criteria and were randomized to a 6-month waitlist (n=372) or the PCHMS (n=370). Amongst the 470 participants eligible for primary analysis, PCHMS users were 6.7% (95% CI: 1.46 to 12.30) more likely than the waitlist to receive an influenza vaccine (waitlist: 4.9% (12/246, 95% CI 2.8 to 8.3) vs PCHMS: 11.6% (26/224, 95% CI 8.0 to 16.5); χ(2)=7.1, p=0.008). PCHMS participants were also 11.6% (95% CI 3.6 to 19.5) more likely to visit the health service provider (waitlist: 17.9% (44/246, 95% CI 13.6 to 23.2) vs PCHMS: 29.5% (66/224, 95% CI: 23.9 to 35.7); χ(2)=8.8, p=0.003). A dose-response effect was detected, where greater use of the PCHMS was associated with higher rates of vaccination (p=0.001) and health service provider visits (p=0.003). DISCUSSION: PCHMS can significantly increase consumer participation in preventive health activities, such as influenza vaccination. CONCLUSIONS: Integrating a PCHMS into routine health service delivery systems appears to be an effective mechanism for enhancing consumer engagement in preventive health measures. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12610000386033. http://www.anzctr.org.au/trial_view.aspx?id=335463.
Annie Y. S. Lau, Vitali Sintchenko, Jacinta Crimmins, Farah Magrabi, Blanca Gallego, Enrico W. Coiera
J. Am. Medical Informatics Assoc.6
2012 A systematic review of the psychological literature on interruption and its patient safety implications
abstract
OBJECTIVE: To understand the complex effects of interruption in healthcare. MATERIALS AND METHODS: As interruptions have been well studied in other domains, the authors undertook a systematic review of experimental studies in psychology and human-computer interaction to identify the task types and variables influencing interruption effects. RESULTS: 63 studies were identified from 812 articles retrieved by systematic searches. On the basis of interruption profiles for generic tasks, it was found that clinical tasks can be distinguished into three broad types: procedural, problem-solving, and decision-making. Twelve experimental variables that influence interruption effects were identified. Of these, six are the most important, based on the number of studies and because of their centrality to interruption effects, including working memory load, interruption position, similarity, modality, handling strategies, and practice effect. The variables are explained by three main theoretical frameworks: the activation-based goal memory model, prospective memory, and multiple resource theory. DISCUSSION: This review provides a useful starting point for a more comprehensive examination of interruptions potentially leading to an improved understanding about the impact of this phenomenon on patient safety and task efficiency. The authors provide some recommendations to counter interruption effects. CONCLUSION: The effects of interruption are the outcome of a complex set of variables and should not be considered as uniformly predictable or bad. The task types, variables, and theories should help us better to identify which clinical tasks and contexts are most susceptible and assist in the design of information systems and processes that are resilient to interruption.
Simon Y. W. Li, Farah Magrabi, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2012 Using FDA reports to inform a classification for health information technology safety problems
abstract
OBJECTIVE: To expand an emerging classification for problems with health information technology (HIT) using reports submitted to the US Food and Drug Administration Manufacturer and User Facility Device Experience (MAUDE) database. DESIGN: HIT events submitted to MAUDE were retrieved using a standardized search strategy. Using an emerging classification with 32 categories of HIT problems, a subset of relevant events were iteratively analyzed to identify new categories. Two coders then independently classified the remaining events into one or more categories. Free-text descriptions were analyzed to identify the consequences of events. MEASUREMENTS: Descriptive statistics by number of reported problems per category and by consequence; inter-rater reliability analysis using the κ statistic for the major categories and consequences. RESULTS: A search of 899 768 reports from January 2008 to July 2010 yielded 1100 reports about HIT. After removing duplicate and unrelated reports, 678 reports describing 436 events remained. The authors identified four new categories to describe problems with software functionality, system configuration, interface with devices, and network configuration; the authors' classification with 32 categories of HIT problems was expanded by the addition of these four categories. Examination of the 436 events revealed 712 problems, 96% were machine-related, and 4% were problems at the human-computer interface. Almost half (46%) of the events related to hazardous circumstances. Of the 46 events (11%) associated with patient harm, four deaths were linked to HIT problems (0.9% of 436 events). CONCLUSIONS: Only 0.1% of the MAUDE reports searched were related to HIT. Nevertheless, Food and Drug Administration reports did prove to be a useful new source of information about the nature of software problems and their safety implications with potential to inform strategies for safe design and implementation.
Farah Magrabi, Mei-Sing Ong, William Runciman, Enrico W. Coiera
J. Am. Medical Informatics Assoc.4
2012 Automated identification of extreme-risk events in clinical incident reports
abstract
OBJECTIVES: To explore the feasibility of using statistical text classification to automatically detect extreme-risk events in clinical incident reports. METHODS: Statistical text classifiers based on Naïve Bayes and Support Vector Machine (SVM) algorithms were trained and tested on clinical incident reports to automatically detect extreme-risk events, defined by incidents that satisfy the criteria of Severity Assessment Code (SAC) level 1. For this purpose, incident reports submitted to the Advanced Incident Management System by public hospitals from one Australian region were used. The classifiers were evaluated on two datasets: (1) a set of reports with diverse incident types (n=120); (2) a set of reports associated with patient misidentification (n=166). Results were assessed using accuracy, precision, recall, F-measure, and area under the curve (AUC) of receiver operating characteristic curves. RESULTS: The classifiers performed well on both datasets. In the multi-type dataset, SVM with a linear kernel performed best, identifying 85.8% of SAC level 1 incidents (precision=0.88, recall=0.83, F-measure=0.86, AUC=0.92). In the patient misidentification dataset, 96.4% of SAC level 1 incidents were detected when SVM with linear, polynomial or radial-basis function kernel was used (precision=0.99, recall=0.94, F-measure=0.96, AUC=0.98). Naïve Bayes showed reasonable performance, detecting 80.8% of SAC level 1 incidents in the multi-type dataset and 89.8% of SAC level 1 patient misidentification incidents. Overall, higher prediction accuracy was attained on the specialized dataset, compared with the multi-type dataset. CONCLUSION: Text classification techniques can be applied effectively to automate the detection of extreme-risk events in clinical incident reports.
Mei-Sing Ong, Farah Magrabi, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2011 Computational inference of grammars for larger-than-gene structures from annotated gene sequences
abstract
MOTIVATION: Larger than gene structures (LGS) are DNA segments that include at least one gene and often other segments such as inverted repeats and gene promoters. Mobile genetic elements (MGE) such as integrons are LGS that play an important role in horizontal gene transfer, primarily in Gram-negative organisms. Known LGS have a profound effect on organism virulence, antibiotic resistance and other properties of the organism due to the number of genes involved. Expert-compiled grammars have been shown to be an effective computational representation of LGS, well suited to automating annotation, and supporting de novo gene discovery. However, development of LGS grammars by experts is labour intensive and restricted to known LGS. OBJECTIVES: This study uses computational grammar inference methods to automate LGS discovery. We compare the ability of six algorithms to infer LGS grammars from DNA sequences annotated with genes and other short sequences. We compared the predictive power of learned grammars against an expert-developed grammar for gene cassette arrays found in Class 1, 2 and 3 integrons, which are modular LGS containing up to 9 of about 240 cassette types. RESULTS: Using a Bayesian generalization algorithm our inferred grammar was able to predict > 95% of MGE structures in a corpus of 1760 sequences obtained from Genbank (F-score 75%). Even with 100% noise added to the training and test sets, we obtained an F-score of 68%, indicating that the method is robust and has the potential to predict de novo LGS structures when the underlying gene features are known. AVAILABILITY: http://www2.chi.unsw.edu.au/attacca.
Guy Tsafnat, Jaron Schaeffer, Andrew Clayphan, Jonathan R. Iredell, Sally R. Partridge, Enrico W. Coiera
Bioinform.6
2011 Agreement between common goals discussed and documented in the ICU
abstract
OBJECTIVE: Meaningful use of electronic health records (EHRs) is dependent on accurate clinical documentation. Documenting common goals in the intensive care unit (ICU), such as sedation and ventilator management plans, may increase collaboration and decrease patient length of stay. This study analyzed the degree to which goals stated were present in the EHR. DESIGN: Descriptive correlational study of common goals verbally stated during daily ICU interdisciplinary rounds compared with the presence of those goals, and actions related to those goals, documented in the EHR over the subsequent 24 h for 28 patients over 15 days. The study setting was a neurovascular ICU with a fully implemented electronic nursing and physician documentation system. MEASUREMENTS: Descriptive statistics and χ(2) analyses were used to assess differences in EHR documentation of stated goals and goal-related actions. Inter-coder reliability was performed on 16 (13%) of the 127 stated goals. RESULTS: One-quarter of the stated goals were not documented in the EHR. If a goal was not documented, actions related to that goal were 60% less likely to be documented. The attending physician note contained 81% of the stated ventilator weaning goals, but only 49% of the sedation weaning goals; additionally, sedation goals were not part of the structured nursing documentation. Inter-coder reliability (κ) was greater than 0.82. LIMITATIONS: Observations in a single ICU setting at a large academic medical center using a commercial EHR. CONCLUSION: The current documentation tools available in EHRs may not be sufficient to capture common goals of ICU patient care.
Sarah A. Collins, Suzanne Bakken, David K. Vawdrey, Enrico W. Coiera, Leanne M. Currie
J. Am. Medical Informatics Assoc.4
2011 A simulation framework for mapping risks in clinical processes: the case of in-patient transfers
abstract
OBJECTIVE: To model how individual violations in routine clinical processes cumulatively contribute to the risk of adverse events in hospital using an agent-based simulation framework. DESIGN: An agent-based simulation was designed to model the cascade of common violations that contribute to the risk of adverse events in routine clinical processes. Clinicians and the information systems that support them were represented as a group of interacting agents using data from direct observations. The model was calibrated using data from 101 patient transfers observed in a hospital and results were validated for one of two scenarios (a misidentification scenario and an infection control scenario). Repeated simulations using the calibrated model were undertaken to create a distribution of possible process outcomes. The likelihood of end-of-chain risk is the main outcome measure, reported for each of the two scenarios. RESULTS: The simulations demonstrate end-of-chain risks of 8% and 24% for the misidentification and infection control scenarios, respectively. Over 95% of the simulations in both scenarios are unique, indicating that the in-patient transfer process diverges from prescribed work practices in a variety of ways. CONCLUSIONS: The simulation allowed us to model the risk of adverse events in a clinical process, by generating the variety of possible work subject to violations, a novel prospective risk analysis method. The in-patient transfer process has a high proportion of unique trajectories, implying that risk mitigation may benefit from focusing on reducing complexity rather than augmenting the process with further rule-based protocols.
Adam G. Dunn, Mei-Sing Ong, Johanna I. Westbrook, Farah Magrabi, Enrico W. Coiera, Wayne Wobcke
J. Am. Medical Informatics Assoc.5
2010 Errors and electronic prescribing: a controlled laboratory study to examine task complexity and interruption effects
abstract
OBJECTIVE: To examine the effect of interruptions and task complexity on error rates when prescribing with computerized provider order entry (CPOE) systems, and to categorize the types of prescribing errors. DESIGN: Two within-subject factors: task complexity (complex vs simple) and interruption (interruption vs no interruption). Thirty-two hospital doctors used a CPOE system in a computer laboratory to complete four prescribing tasks, half of which were interrupted using a counterbalanced design. MEASUREMENTS: Types of prescribing errors, error rate, resumption lag, and task completion time. RESULTS: Errors in creating and updating electronic medication charts that were measured included failure to enter allergy information; selection of incorrect medication, dose, route, formulation, or frequency of administration from lists and drop-down menus presented by the CPOE system; incorrect entry or omission in entering administration times, start date, and free-text qualifiers; and omissions in prescribing and ceasing medications. When errors occurred, the error rates across the four prescribing tasks ranged from 0.5% (1 incorrect medication selected out of 192 chances for selecting a medication or error opportunities) to 16% (5 failures to enter allergy information out of 32 error opportunities). Any impact of interruptions on prescribing error rates and task completion times was not detected in our experiment. However, complex tasks took significantly longer to complete (F(1, 27)=137.9; p<0.001) and when execution was interrupted they required almost three times longer to resume compared to simple tasks (resumption lag complex=9.6 seconds, SD=5.6; resumption lag simple=3.4 seconds, SD=1.7; t(28)=6.186; p<0.001). CONCLUSION: Most electronic prescribing errors found in this study could be described as slips in using the CPOE system to create and update electronic medication charts. Cues available within the user interface may have aided resumption of interrupted tasks making CPOE systems robust to some interruption effects. Further experiments are required to rule out any effect interruption might have on CPOE error rates.
Farah Magrabi, Simon Y. W. Li, Richard O. Day, Enrico W. Coiera
J. Am. Medical Informatics Assoc.4
2010 An analysis of computer-related patient safety incidents to inform the development of a classification
abstract
OBJECTIVE: To analyze patient safety incidents associated with computer use to develop the basis for a classification of problems reported by health professionals. DESIGN: Incidents submitted to a voluntary incident reporting database across one Australian state were retrieved and a subset (25%) was analyzed to identify 'natural categories' for classification. Two coders independently classified the remaining incidents into one or more categories. Free text descriptions were analyzed to identify contributing factors. Where available medical specialty, time of day and consequences were examined. MEASUREMENTS: Descriptive statistics; inter-rater reliability. RESULTS: A search of 42,616 incidents from 2003 to 2005 yielded 123 computer related incidents. After removing duplicate and unrelated incidents, 99 incidents describing 117 problems remained. A classification with 32 types of computer use problems was developed. Problems were grouped into information input (31%), transfer (20%), output (20%) and general technical (24%). Overall, 55% of problems were machine related and 45% were attributed to human-computer interaction. Delays in initiating and completing clinical tasks were a major consequence of machine related problems (70%) whereas rework was a major consequence of human-computer interaction problems (78%). While 38% (n=26) of the incidents were reported to have a noticeable consequence but no harm, 34% (n=23) had no noticeable consequence. CONCLUSION: Only 0.2% of all incidents reported were computer related. Further work is required to expand our classification using incident reports and other sources of information about healthcare IT problems. Evidence based user interface design must focus on the safe entry and retrieval of clinical information and support users in detecting and correcting errors and malfunctions.
Farah Magrabi, Mei-Sing Ong, William Runciman, Enrico W. Coiera
J. Am. Medical Informatics Assoc.4
2009 In silico prioritisation of candidate genes for prokaryotic gene function discovery: an application of phylogenetic profiles
abstract
BACKGROUND: In silico candidate gene prioritisation (CGP) aids the discovery of gene functions by ranking genes according to an objective relevance score. While several CGP methods have been described for identifying human disease genes, corresponding methods for prokaryotic gene function discovery are lacking. Here we present two prokaryotic CGP methods, based on phylogenetic profiles, to assist with this task. RESULTS: Using gene occurrence patterns in sample genomes, we developed two CGP methods (statistical and inductive CGP) to assist with the discovery of bacterial gene functions. Statistical CGP exploits the differences in gene frequency against phenotypic groups, while inductive CGP applies supervised machine learning to identify gene occurrence pattern across genomes. Three rediscovery experiments were designed to evaluate the CGP frameworks. The first experiment attempted to rediscover peptidoglycan genes with 417 published genome sequences. Both CGP methods achieved best areas under receiver operating characteristic curve (AUC) of 0.911 in Escherichia coli K-12 (EC-K12) and 0.978 Streptococcus agalactiae 2603 (SA-2603) genomes, with an average improvement in precision of >3.2-fold and a maximum of >27-fold using statistical CGP. A median AUC of >0.95 could still be achieved with as few as 10 genome examples in each group of genome examples in the rediscovery of the peptidoglycan metabolism genes. In the second experiment, a maximum of 109-fold improvement in precision was achieved in the rediscovery of anaerobic fermentation genes in EC-K12. The last experiment attempted to rediscover genes from 31 metabolic pathways in SA-2603, where 14 pathways achieved AUC >0.9 and 28 pathways achieved AUC >0.8 with the best inductive CGP algorithms. CONCLUSION: Our results demonstrate that the two CGP methods can assist with the study of functionally uncategorised genomic regions and discovery of bacterial gene-function relationships. Our rediscovery experiments also provide a set of standard tasks against which future methods may be compared.
Frank P. Y. Lin, Enrico W. Coiera, Ruiting Lan, Vitali Sintchenko
BMC Bioinform.2
2009 Towards bioinformatics assisted infectious disease control
abstract
BACKGROUND: This paper proposes a novel framework for bioinformatics assisted biosurveillance and early warning to address the inefficiencies in traditional surveillance as well as the need for more timely and comprehensive infection monitoring and control. It leverages on breakthroughs in rapid, high-throughput molecular profiling of microorganisms and text mining. RESULTS: This framework combines the genetic and geographic data of a pathogen to reconstruct its history and to identify the migration routes through which the strains spread regionally and internationally. A pilot study of Salmonella typhimurium genotype clustering and temporospatial outbreak analysis demonstrated better discrimination power than traditional phage typing. Half of the outbreaks were detected in the first half of their duration. CONCLUSION: The microbial profiling and biosurveillance focused text mining tools can enable integrated infectious disease outbreak detection and response environments based upon bioinformatics knowledge models and measured by outcomes including the accuracy and timeliness of outbreak detection.
Vitali Sintchenko, Blanca Gallego, Grace Chung, Enrico W. Coiera
BMC Bioinform.4
2009 Context-driven discovery of gene cassettes in mobile integrons using a computational grammar
abstract
BACKGROUND: Gene discovery algorithms typically examine sequence data for low level patterns. A novel method to computationally discover higher order DNA structures is presented, using a context sensitive grammar. The algorithm was applied to the discovery of gene cassettes associated with integrons. The discovery and annotation of antibiotic resistance genes in such cassettes is essential for effective monitoring of antibiotic resistance patterns and formulation of public health antibiotic prescription policies. RESULTS: We discovered two new putative gene cassettes using the method, from 276 integron features and 978 GenBank sequences. The system achieved kappa = 0.972 annotation agreement with an expert gold standard of 300 sequences. In rediscovery experiments, we deleted 789,196 cassette instances over 2030 experiments and correctly relabelled 85.6% (alpha > or = 95%, E < or = 1%, mean sensitivity = 0.86, specificity = 1, F-score = 0.93), with no false positives.Error analysis demonstrated that for 72,338 missed deletions, two adjacent deleted cassettes were labeled as a single cassette, increasing performance to 94.8% (mean sensitivity = 0.92, specificity = 1, F-score = 0.96). CONCLUSION: Using grammars we were able to represent heuristic background knowledge about large and complex structures in DNA. Importantly, we were also able to use the context embedded in the model to discover new putative antibiotic resistance gene cassettes. The method is complementary to existing automatic annotation systems which operate at the sequence level.
Guy Tsafnat, Enrico W. Coiera, Sally R. Partridge, Jaron Schaeffer, Jonathan R. Iredell
BMC Bioinform.2
2009 Research Paper: Building a National Health IT System from the Middle Out
abstract
The top-down approach of many national programs for healthcare information technology (IT) may be at the heart of their current problems. The medical-industrial complex loves a big procurement, and the contracts do not get much bigger than for building nation-scale health information systems (NHIS). But do we really need government embedded in the process of IT implementation, something it so clearly and routinely struggles with? Or is it better for government to simply set the policy rules of the game, given that it is policy in which they are expert? As the new United States Administration has recently signalled a massive injection of funds into building a National Health Information infrastructure via the American Recovery and Reinvestment Act (ARRA), what lessons can be learned from the past, and what strategic shape should the Federal intervention take? The English National Health System (NHS) National Program for IT (NPfIT) in many ways serves as an international beacon for healthcare reform, because of its clear message that major restructuring of health services is not possible without a pervasive information infrastructure. The NPfIT is rolling out working systems and delivering tangible benefits to patients and caregivers. Yet no one could deny that there have been plenty of setbacks, misgivings, clinical unrest, delays, cost overruns, and paring back of promised functionality, culminating in demands from some political quarters to shut down the program.1 The NPfIT was bound to experience some difficulties purely on the basis of its scale and complexity.2 However, it is becoming apparent that there may be another, more foundational, cause of NPfIT's problems. The NHS remains one of the few nation-scale, single-payer health systems in the world. It thus has nation-scale management and governance structures to match, and these inevitably encourage a top-down system architecture, standards compliance, and procurement process. …
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
2009 Research Paper: Can Cognitive Biases during Consumer Health Information Searches Be Reduced to Improve Decision Making?
abstract
OBJECTIVE: To test whether the anchoring and order cognitive biases experienced during search by consumers using information retrieval systems can be corrected to improve the accuracy of, and confidence in, answers to health-related questions. DESIGN: A prospective study was conducted on 227 undergraduate students who used an online search engine developed by the authors to find health information and then answer six randomly assigned consumer health questions. The search engine was fitted with a baseline user interface and two modified interfaces specifically designed to debias anchoring or order effect. Each subject used all three user interfaces, answering two questions with each. MEASUREMENTS: Frequencies of correct answers pre- and post- search and confidence in answers were collected. Time taken to search and then answer a question, the number of searches conducted and the number of links accessed in a search session were also recorded. User preferences for each interface were measured. Chi-square analyses tested for the presence of biases with each user interface. The Kolmogorov-Smirnov test checked for equality of distribution of the evidence analyzed for each user interface. The test for difference between proportions and the Wilcoxon signed ranks test were used when comparing interfaces. RESULTS: Anchoring and order effects were present amongst subjects using the baseline search interface (anchoring: p < 0.001; order: p = 0.026). With use of the order debiasing interface, the initial order effect was no longer present (p = 0.34) but there was no significant improvement in decision accuracy (p = 0.23). While the anchoring effect persisted when using the anchor debiasing interface (p < 0.001), its use was associated with a 10.3% increase in subjects who had answered incorrectly pre-search, answering correctly post-search (p = 0.10). Subjects using either debiasing user interface conducted fewer searches and accessed more documents compared to baseline (p < 0.001). In addition, the majority of subjects preferred using a debiasing interface over baseline. CONCLUSION: This study provides evidence that (i) debiasing strategies can be integrated into the user interface of a search engine; (ii) information interpretation behaviors can be to some extent debiased; and that (iii) attempts to debias information searching by consumers can influence their ability to answer health-related questions accurately, their confidence in these answers, as well as the strategies used to conduct searches and retrieve information.
Annie Y. S. Lau, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
2009 Viewpoint Paper: Computational Reasoning across Multiple Models
abstract
Computational support of clinical decisions frequently requires the integration of data in a variety of formats and from multiple sources and domains. Some impressive multiscale computational models of biological phenomena have been developed as part of the study of disease and healthcare systems. One can now contemplate harnessing these models arising from computational biology and using highly interconnected clinical data to support clinical decision-making. Indeed, understanding how to build computational systems able to reason across heterogeneous models and datasets is one of the major and perhaps foundational challenges of translational biomedical informatics. In this paper, the authors examine the use of multimodels (models composed of several daughter models) and explore three major research challenges to reasoning across multiple models: model selection, model composition, and computer aided model construction.
Guy Tsafnat, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
2009 Biosurveillance of emerging biothreats using scalable genotype clustering
Blanca Gallego, Vitali Sintchenko, Qinning Wang, Lester Hiley, Gwendolyn L. Gilbert, Enrico W. Coiera
J. Biomed. Informatics6
2008 Research Paper: Is Relevance Relevant? User Relevance Ratings May Not Predict the Impact of Internet Search on Decision Outcomes
abstract
OBJECTIVE: A common measure of Internet search engine effectiveness is its ability to find documents that a user perceives as 'relevant'. This study sought to test whether user provided relevance ratings for documents retrieved by an Internet search engine correlate with the decision outcome after use of a search engine. DESIGN: 227 university students were asked to answer four randomly assigned consumer health questions, then to conduct an Internet search on one of two randomly assigned search engines of different performance, and to again answer the question. MEASUREMENTS: Participants were asked to provide a relevance score for each document retrieved as well as a pre and post search answer to each question. RESULTS: User relevance rankings had little or no predictive power. Relevance rankings were unable to predict whether the user of a search engine could correctly answer a question after search and could not differentiate between two search engines with statistically different performance in the hands of users. Only when users had strong prior knowledge of the questions, and the decision task was of low complexity, did relevance appear to have modest predictive power. CONCLUSIONS: User provided relevance rankings taken in isolation seem to be of limited to no value when designing a search engine that will be used in a general-purpose setting. Relevance rankings may have a place in situations in which experts provide rankings, and decision tasks are of complexity commensurate with the abilities of the raters. A more natural metric of search engine performance may be a user's ability to accurately complete a task, as this removes the inherent subjectivity of relevance rankings, and provides a direct and repeatable outcome measure which directly correlates with the performance of the search technology in the hands of users.
Enrico W. Coiera, Victor Vickland
J. Am. Medical Informatics Assoc.1
2008 Research Paper: Clinical Decision Velocity is Increased when Meta-search Filters Enhance an Evidence Retrieval System
abstract
OBJECTIVE: To test whether the use of an evidence retrieval system that uses clinically targeted meta-search filters can enhance the rate at which clinicians make correct decisions, reduce the effort involved in locating evidence, and provide an intuitive match between clinical tasks and search filters. DESIGN: A laboratory experiment under controlled conditions asked 75 clinicians to answer eight randomly sequenced clinical questions, using one of two randomly assigned search engines. The first search engine Quick Clinical (QC) was equipped with meta-search filters (the combined use of meta-search and search filters) designed to answer typical clinical questions e.g., treatment, diagnosis, and the second 'library model' system (LM) offered free access to an identical evidence set with no filter support. MEASUREMENTS: Changes in clinical decision making were measured by the proportion of correct post-search answers provided to questions, the time taken to answer questions, and the number of searches and links to documents followed in a search session. The intuitive match between meta-search filters and clinical tasks was measured by the proportion and distribution of filters selected for individual clinical questions. RESULTS: Clinicians in the two groups performed equally well pre-search. Post search answers improved overall by 21%, with 52.2% of answers correct with QC and 54.7% with LM (chi(2) = 0.33, df = 1, p > 0.05). Users of QC obtained a significantly greater percentage of their correct answers within the first two minutes of searching compared to LM users (QC 58.2%; LM 32.9%; chi(2) = 19.203, df = 1, p < 0.001). There was a statistical difference for QC and LM survival curves, which plotted overall time to answer questions, irrespective of answer (Wilcoxon, p = 0.019) and for the average time to provide a correct answer (Wilcoxon, p = 0.006). The QC system users conducted significantly fewer searches per scenario (m = 3.0 SD = 1.15 versus m = 5.5 SD1.97, t = 6.63, df = 72, p = 0.0001). Clinicians using the QC system followed fewer document links than did those who used LM (respectively 3.9 links SD = 1.20 versus 4.7 links SD = 1.79, t = 2.13, df = 72, p = 0.0368). In 6 of the 8 questions, two meta-search filters accounted for 89% or more of clinicians' first choice, suggesting the choice of filter intuitively matched the clinical decision task at hand. CONCLUSIONS: Meta-search filters result in clinicians arriving at answers more quickly than unconstrained searches across information sources, and appear to increase the rate with which correct decisions are made. In time restricted clinical settings meta-search filters may thus improve overall decision accuracy, as fewer searches that could otherwise lead to a correct answer are abandoned. Meta-search filters appear to be intuitive to use, suggesting that the simplicity of the user model would fit very well into clinical settings.
Enrico W. Coiera, Johanna I. Westbrook, Kris Rogers
J. Am. Medical Informatics Assoc.1
2007 Research Paper: Do People Experience Cognitive Biases while Searching for Information?
abstract
OBJECTIVE: To test whether individuals experience cognitive biases whilst searching using information retrieval systems. Biases investigated are anchoring, order, exposure and reinforcement. DESIGN: A retrospective analysis and a prospective experiment were conducted to investigate whether cognitive biases affect the way that documentary evidence is interpreted while searching online. The retrospective analysis was conducted on the search and decision behaviors of 75 clinicians (44 doctors, 31 nurses), answering questions for 8 clinical scenarios within 80 minutes in a controlled setting. The prospective study was conducted on 227 undergraduate students, who used the same search engine to answer two of six randomly assigned consumer health questions. MEASUREMENTS: Frequencies of correct answers pre- and post- search, and confidence in answers were collected. The impact of reading a document on the final decision was measured by the population likelihood ratio (LR) of the frequency of reading the document and the frequency of obtaining a correct answer. Documents with a LR > 1 were most likely to be associated with a correct answer, and those with a LR < 1 were most likely to be associated with an incorrect answer to a question. Agreement between a subject and the evidence they read was estimated by a concurrence rate, which measured the frequency that subjects' answers agreed with the likelihood ratios of a group of documents, normalized for document order, time exposure or reinforcement through repeated access. Serial position curves were plotted for the relationship between subjects' pre-search confidence, document order, the number of times and length of time a document was accessed, and concurrence with post-search answers. Chi-square analyses tested for the presence of biases, and the Kolmogorov-Smirnov test checked for equality of distribution of evidence in the comparison populations. RESULTS: A person's prior belief (anchoring) has a significant impact on their post-search answer (retrospective: P < 0.001; prospective: P < 0.001). Documents accessed at different positions in a search session (order effect [retrospective: P = 0.76; prospective: P = 0.026]), and documents processed for different lengths of time (exposure effect [retrospective: P = 0.27; prospective: P = 0.0081]) also influenced decision post-search more than expected in the prospective experiment but not in the retrospective analysis. Reinforcement through repeated exposure to a document did not yield statistical differences in decision outcome post-search (retrospective: P = 0.31; prospective: P = 0.81). CONCLUSION: People may experience anchoring, exposure and order biases while searching for information, and these biases may influence the quality of decision making during and after the use of information retrieval systems.
Annie Y. S. Lau, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
2007 Model Formulation: Multimethod Evaluation of Information and Communication Technologies in Health in the Context of Wicked Problems and Sociotechnical Theory
abstract
OBJECTIVE: Few research designs look at the deep structure of complex social systems. We report the design and implementation of a multimethod evaluation model to assess the impact of computerized order entry systems on both the technical and social systems within a health care organization. DESIGN: We designed a multimethod evaluation model informed by sociotechnical theory and an appreciation of the nature of wicked problems. We mobilized this model to assess the impact of an electronic medication management system via a three-year program of research at a major academic hospital. MEASUREMENTS: Model components include measurements relating to three dimensions of system impact: safety and quality, organizational culture, and work and communication patterns. RESULTS: Application of the evaluation model required the development and testing of purpose-built measurement tools such as software to collect multidimensional work measurement data. The model applied established research methods including medication error audits and social network analysis. Design features of these tools and techniques are described, along with the practical challenges of their implementation. The distinctiveness of doing research within a unique paradigm of complex systems, explicating the wickedness and the dimensionality of sociotechnical theory, is articulated. CONCLUSION: Designing an effective evaluation model requires a deep understanding of the nature and complexity of the problems that information technology interventions in health care are trying to address. Adopting a sociotechnical perspective for model generation improves our ability to develop evaluation models that are adaptive and sensitive to the characteristics of wicked problems and provides a strong theoretical basis from which to analyze and interpret findings.
Johanna I. Westbrook, Jeffrey Braithwaite, Andrew Georgiou, Amanda Ampt, Nerida Creswick, Enrico W. Coiera, Rick Iedema
J. Am. Medical Informatics Assoc.6
2006 Decision Complexity Affects the Extent and Type of Decision Support Use
Vitali Sintchenko, Enrico W. Coiera
AMIA2
2006 A Bayesian model that predicts the impact of Web searching on decision making
abstract
Abstract This study aimed to develop a model for predicting the impact of information access using Web searches, on human decision making. Models were constructed using a database of search behaviors and decisions of 75 clinicians, who answered questions about eight scenarios within 80 minutes in a controlled setting at a university computer laboratory. Bayesian models were developed with and without bias factors to account for anchoring, primacy, recency, exposure, and reinforcement decision biases. Prior probabilities were estimated from the population prior, from a personal prior calculated from presearch answers and confidence ratings provided by the participants, from an overall measure of willingness to switch belief before and after searching, and from a willingness to switch belief calculated in each individual scenario. The optimal Bayes model predicted user answers in 73.3% (95% CI: 68.71 to 77.35%) of cases, and incorporated participants' willingness to switch belief before and after searching for each scenario, as well as the decision biases they encounter during the search journey. In most cases, it is possible to predict the impact of a sequence of documents retrieved by a Web search engine on a decision task without reference to the content or structure of the documents, but relying solely on a simple Bayesian model of belief revision.
Annie Y. S. Lau, Enrico W. Coiera
J. Assoc. Inf. Sci. Technol.2
2005 Application of Information Technology: Handheld Computer-based Decision Support Reduces Patient Length of Stay and Antibiotic Prescribing in Critical Care
abstract
OBJECTIVE: This study assessed the effect of a handheld computer-based decision support system (DSS) on antibiotic use and patient outcomes in a critical care unit. DESIGN: A DSS containing four types of evidence (patient microbiology reports, local antibiotic guidelines, unit-specific antibiotic susceptibility data for common bacterial pathogens, and a clinical pulmonary infection score calculator) was developed and implemented on a handheld computer for use in the intensive care unit at a tertiary referral hospital. System impact was assessed in a prospective "before/after" cohort trial lasting 12 months. Outcome measures were defined daily doses (DDDs) of antibiotics per 1,000 patient-days, patient length of stay, and mortality. RESULTS: The number of admissions, APACHE (Acute Physiology, Age, and Chronic Health Evaluation) II and SAPS (Simplified Acute Physiology Score) II for patients in preintervention, and intervention (DSS use) periods were statistically comparable. The mean patient length of stay and the use of antibiotics in the unit during six months of the DSS use decreased from 7.15 to 6.22 bed-days (p = 0.02) and from 1,767 DDD to 1,458 DDD per 1,000 patient-days (p = 0.04), respectively, with no change in mortality. The DSS was accessed 674 times during 168 days of the trial. Microbiology reports and antibiotic guidelines were the two most commonly used (53% and 22.5%, respectively) types of evidence. The greatest reduction was observed in the use of beta-lactamase-resistant penicillins and vancomycin. CONCLUSION: Handheld computer-based decision support contributed to a significant reduction in patient length of stay and antibiotic prescribing in a critical care unit.
Vitali Sintchenko, Jonathan R. Iredell, Gwendolyn L. Gilbert, Enrico W. Coiera
J. Am. Medical Informatics Assoc.4
2005 Research Paper: Do Online Information Retrieval Systems Help Experienced Clinicians Answer Clinical Questions?
abstract
OBJECTIVE: To assess the impact of clinicians' use of an online information retrieval system on their performance in answering clinical questions. DESIGN: Pre-/post-intervention experimental design. MEASUREMENTS: In a computer laboratory, 75 clinicians (26 hospital-based doctors, 18 family practitioners, and 31 clinical nurse consultants) provided 600 answers to eight clinical scenarios before and after the use of an online information retrieval system. We examined the proportion of correct answers pre- and post-intervention, direction of change in answers, and differences between professional groups. RESULTS: System use resulted in a 21% improvement in clinicians' answers, from 29% (95% confidence interval [CI] 25.4-32.6) correct pre- to 50% (95% CI 46.0-54.0) post-system use. In 33% (95% CI 29.1-36.9) answers were changed from incorrect to correct. In 21% (95% CI 17.1-23.9) correct pre-test answers were supported by evidence found using the system, and in 7% (95% CI 4.9-9.1) correct pre-test answers were changed incorrectly. For 40% (35.4-43.6) of scenarios, incorrect pre-test answers were not rectified following system use. Despite significant differences in professional groups' pre-test scores [family practitioners: 41% (95% CI 33.0-49.0), hospital doctors: 35% (95% CI 28.5-41.2), and clinical nurse consultants: 17% (95% CI 12.3-21.7; chi(2) = 29.0, df = 2, p < 0.01)], there was no difference in post-test scores. (chi(2) = 2.6, df = 2, p = 0.73). CONCLUSIONS: The use of an online information retrieval system was associated with a significant improvement in the quality of answers provided by clinicians to typical clinical problems. In a small proportion of cases, use of the system produced errors. While there was variation in the performance of clinical groups when answering questions unaided, performance did not differ significantly following system use. Online information retrieval systems can be an effective tool in improving the accuracy of clinicians' answers to clinical questions.
Johanna I. Westbrook, Enrico W. Coiera, A. Sophie Gosling
J. Am. Medical Informatics Assoc.2
2004 Viewpoint Paper: Some Unintended Consequences of Information Technology in Health Care: The Nature of Patient Care Information System-related Errors
abstract
Medical error reduction is an international issue, as is the implementation of patient care information systems (PCISs) as a potential means to achieving it. As researchers conducting separate studies in the United States, The Netherlands, and Australia, using similar qualitative methods to investigate implementing PCISs, the authors have encountered many instances in which PCIS applications seem to foster errors rather than reduce their likelihood. The authors describe the kinds of silent errors they have witnessed and, from their different social science perspectives (information science, sociology, and cognitive science), they interpret the nature of these errors. The errors fall into two main categories: those in the process of entering and retrieving information, and those in the communication and coordination process that the PCIS is supposed to support. The authors believe that with a heightened awareness of these issues, informaticians can educate, design systems, implement, and conduct research in such a way that they might be able to avoid the unintended consequences of these subtle silent errors.
Joan S. Ash, Marc Berg, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2004 Viewpoint Paper: e-Consent: The Design And Implementation of Consumer Consent Mechanisms in an Electronic Environment
abstract
The effective coordination of health care relies on communication of confidential information about consumers between different health and community care services. However, consumers must be able to give or withhold "e-Consent" to those who wish to access their electronic health information. There are several possible forms for e-Consent. In the general consent model, a patient provides blanket consent for access to his or her information by an organization for all future information requests. Conversely, general denial explicitly denies consent for information to be used in future circumstances, and in each new episode of care, a new consent would be needed to obtain information. In the general consent with specific denial model, a patient attaches specific exclusion conditions to his or her general approval to future accesses. In contrast, in the general denial with explicit consent model, a patient issues a blanket block on all future accesses but allows the inclusion of future use under specified conditions. There also are several alternative functions for an e-Consent system. Consent could be captured as a matter of legal record. E-Consent systems could be more active by prompting clinicians to indicate that they have noted consent conditions before they access a record. Finally, the record of patient consent could be fully active and used as a gatekeeper in a distributed information environment. There probably will need to be some form of data object that is associated with patient information. This e-Consent object (or e-Co) will contain the specific conditions under which the data to which it is attached can be retrieved. Given the complexity of clinical work and the substantial variation we can expect in an individual's desire to make his or her personal medical details available, it is unlikely a "one size fits all" approach to e-Consent will work. Consequently, with a well-chosen consent design, it should be possible to balance the specific need for privacy of some of the population against the desire by others to err on the side of clinical safety, and clinicians desire to minimize the burden that an electronic consent mechanism would impose.
Enrico W. Coiera, Roger Clarke
J. Am. Medical Informatics Assoc.1
2004 Research Paper: Comparative Impact of Guidelines, Clinical Data, and Decision Support on Prescribing Decisions: An Interactive Web Experiment with Simulated Cases
abstract
OBJECTIVE: The aim of this study was to compare the clinical impact of computerized decision support with and without electronic access to clinical guidelines and laboratory data on antibiotic prescribing decisions. DESIGN: A crossover trial was conducted of four levels of computerized decision support-no support, antibiotic guidelines, laboratory reports, and laboratory reports plus a decision support system (DSS), randomly allocated to eight simulated clinical cases accessed by the Web. MEASUREMENTS: Rate of intervention adoption was measured by frequency of accessing information support, cost of use was measured by time taken to complete each case, and effectiveness of decision was measured by correctness of and self-reported confidence in individual prescribing decisions. Clinical impact score was measured by adoption rate and decision effectiveness. RESULTS: Thirty-one intensive care and infectious disease specialist physicians (ICPs and IDPs) participated in the study. Ventilator-associated pneumonia treatment guidelines were used in 24 (39%) of the 62 case scenarios for which they were available, microbiology reports in 36 (58%), and the DSS in 37 (60%). The use of all forms of information support did not affect clinicians' confidence in their decisions. Their use of the DSS plus microbiology report improved the agreement of decisions with those of an expert panel from 65% to 97% (p=0.0002), or to 67% (p=0.002) when antibiotic guidelines only were accessed. Significantly fewer IDPs than ICPs accessed information support in making treatment decisions. On average, it took 245 seconds to make a decision using the DSS compared with 113 seconds for unaided prescribing (p<0.001). The DSS plus microbiology reports had the highest clinical impact score (0.58), greater than that of electronic guidelines (0.26) and electronic laboratory reports (0.45). CONCLUSION: When used, computer-based decision support significantly improved decision quality. In measuring the impact of decision support systems, both their effectiveness in improving decisions and their likely rate of adoption in the clinical environment need to be considered. Clinicians chose to use antibiotic guidelines for one third and microbiology reports or the DSS for about two thirds of cases when they were available to assist their prescribing decisions.
Vitali Sintchenko, Enrico W. Coiera, Jonathan R. Iredell, Gwendolyn L. Gilbert
J. Am. Medical Informatics Assoc.2
2004 Research Paper: Do clinicians use online evidence to support patient care? a study of 55, 000 clinicians
abstract
OBJECTIVES: To determine clinicians' (doctors', nurses', and allied health professionals') "actual" and "reported" use of a point-of-care online information retrieval system; and to make an assessment of the extent to which use is related to direct patient care by testing two hypotheses: hypothesis 1: clinicians use online evidence primarily to support clinical decisions relating to direct patient care; and hypothesis 2: clinicians use online evidence predominantly for research and continuing education. DESIGN: Web-log analysis of the Clinical Information Access Program (CIAP), an online, 24-hour, point-of-care information retrieval system available to 55,000 clinicians in public hospitals in New South Wales, Australia. A statewide mail survey of 5,511 clinicians. MEASUREMENTS: Rates of online evidence searching per 100 clinicians for the state and for the 81 individual hospitals studied; reported use of CIAP by clinicians through a self-administered questionnaire; and correlations between evidence searches and patient admissions. RESULTS: Monthly rates of 48.5 "search sessions" per 100 clinicians and 231.6 text hits to single-source databases per 100 clinicians (n = 619,545); 63% of clinicians reported that they were aware of CIAP and 75% of those had used it. Eighty-eight percent of users reported CIAP had the potential to improve patient care and 41% reported direct experience of this. Clinicians' use of CIAP on each day of the week was highly positively correlated with patient admissions (r = 0.99, p < 0.001). This was also true for all ten randomly selected hospitals. CONCLUSION: Clinicians' online evidence use increases with patient admissions, supporting the hypothesis that clinicians' use of evidence is related to direct patient care. Patterns of evidence use and clinicians' self-reports also support this hypothesis.
Johanna I. Westbrook, A. Sophie Gosling, Enrico W. Coiera
J. Am. Medical Informatics Assoc.3
2001 Mediated Agent Interaction
Enrico W. Coiera
AIME1
2000 Viewpoint: Information Economics and the Internet
abstract
Information economics offers insights into the dynamics of information across networked systems like the Internet. An information marketplace is different from other marketplaces because an information good is not actually consumed and can be reproduced and distributed at almost no cost. For information producers to remain profitable, they will need to minimize their exposure to competition. For example, information can be sold by charging site access rather than information access fees, or it can be bundled with other information or "versioned." For information consumers, a variation of Malthus' law predicts that the exponential growth in information will mean that specific information will become increasingly expensive to find, because search costs will grow but human attention will remain limited. Furthermore, the low cost of creating poor-quality information on the Web means that the low-quality information may eventually swamp high-quality resources. The use of reputable information portals on the Web, or smart search technologies, may help in the short run, but it is unclear whether an "information famine" is avoidable in the longer term.
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
2000 Viewpoint: When Conversation Is Better Than Computation
abstract
While largely ignored in informatics thinking, the clinical communication space accounts for the major part of the information flow in health care. Growing evidence indicates that errors in communication give rise to substantial clinical morbidity and mortality. This paper explores the implications of acknowledging the primacy of the communication space in informatics and explores some solutions to communication difficulties. It also examines whether understanding the dynamics of communication between human beings can also improve the way we design information systems in health care. Using the concept of common ground in conversation, proposals are suggested for modeling the common ground between a system and human users. Such models provide insights into when communication or computational systems are better suited to solving information problems.
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
2000 Viewpoint: Improving Clinical Communication: A View from Psychology
abstract
Recent research has studied the communication behaviors of clinical hospital workers and observed a tendency for these workers to use communication behaviors that were often inefficient. Workers were observed to favor synchronous forms of communication, such as telephone calls and chance face-to-face meetings with colleagues, even when these channels were not effective. Synchronous communication also contributes to a highly interruptive working environment, increasing the potential for clinical errors to be made. This paper reviews these findings from a cognitive psychological perspective, focusing on current understandings of how human memory functions and on the potential consequences of interruptions on the ability to work effectively. It concludes by discussing possible communication technology interventions that could be introduced to improve the clinical communication environment and suggests directions for future research.
Julie Parker, Enrico W. Coiera
J. Am. Medical Informatics Assoc.2
1997 Learning Qualitative Models of Dynamic Systems
David T. Hau, Enrico W. Coiera
Mach. Learn.2
1996 Viewpoint: Artificial Intelligence in Medicine: The Challenges Ahead
abstract
The modern study of artificial intelligence in medicine (AIM) is 25 years old. Throughout this period, the field has attracted many of the best computer scientists, and their work represents a remarkable achievement. However, AIM has not been successful-if success is judged as making an impact on the practice of medicine. Much recent work in AIM has been focused inward, addressing problems that are at the crossroads of the parent disciplines of medicine and artificial intelligence. Now, AIM must move forward with the insights that it has gained and focus on finding solutions for problems at the heart of medical practice. The growing emphasis within medicine on evidence-based practice should provide the right environment for that change.
Enrico W. Coiera
J. Am. Medical Informatics Assoc.1
1993 Intelligent monitoring and control of dynamic physiological systems
Enrico W. Coiera
Artif. Intell. Medicine1
1992 Qualitative Superposition
Enrico W. Coiera
Artif. Intell.1
1992 Intermediate depth representations
Enrico W. Coiera
Artif. Intell. Medicine1
1991 The role of domain models in maintaining consistency of large medical knowledge bases
Andrzej J. Glowinski, Enrico W. Coiera, Mike O'Neil
AIME2
1990 Monitoring diseases with empirical and model-generated histories
Enrico W. Coiera
Artif. Intell. Medicine1