VLDB 2026 Research / reviewers in the wild / expert
Farah Magrabi
dblp:45/4277
· DBLP profile ↗
33ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-8426-5588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 5 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-tuning and evaluating large language models for patient safety tasks: classification of contributing factors in incident reportsabstractOBJECTIVE: To evaluate and compare the performance of large language models (LLMs) in identifying contributing factors (CFs) underlying patient safety incident investigations. MATERIALS AND METHODS: Four open-source, lightweight LLMs, including BERT, LLaMA2, GPT2, and Phi-2 were applied to classify CFs across 6 sociotechnical system-levels encompassing 12 categories (eg, person, task, and organizational factors). Reports of real-world patient safety investigations from public health systems were extracted and labelled by domain experts (n_report/CFs = 300/1338). Data were split into training (n = 852), validation (n = 98), and test sets (n = 388). Performance was evaluated using specificity, precision, recall, and F1 scores. RESULTS: The fine-tuned encoder-based BERT model achieved the highest performance, with a micro-averaged F1 score of 63.6%, outperforming all decoder-based models. Among the decoder models, Phi-2 demonstrated the strongest performance (F1 = 54.9%), exceeding both LLaMA2 and GPT2. BERT performed consistently across 6 system-levels but often misclassified "organization" as "person". DISCUSSION: LLMs hold promise for automating the extraction of CFs from complex safety narratives, particularly for frequently reported system-levels such as "person" and "tasks". Such automation may substantially reduce the manual effort required to analyse reports of patient safety investigations while supporting more consistent analysis across large incident datasets. CONCLUSION: Applying LLMs to analyse the underlying causes of patient safety incidents depends on developing high-quality, domain-specific datasets that enhance the representation of patient safety knowledge and improve model understanding of incident causation. Improving data coverage for rare system-levels is essential to address the current limitations of LLMs in capturing nuanced patient safety concepts and domain-specific reasoning. Ying Wang 0003, Lorelle Bowditch, Charlotte Molloy, Yinghua Yu, Peter Hibbert, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 6 |
| 2023 | More than algorithms: an analysis of safety events involving ML-enabled medical devices reported to the FDAabstractOBJECTIVE: To examine the real-world safety problems involving machine learning (ML)-enabled medical devices. MATERIALS AND METHODS: We analyzed 266 safety events involving approved ML medical devices reported to the US FDA's MAUDE program between 2015 and October 2021. Events were reviewed against an existing framework for safety problems with Health IT to identify whether a reported problem was due to the ML device (device problem) or its use, and key contributors to the problem. Consequences of events were also classified. RESULTS: Events described hazards with potential to harm (66%), actual harm (16%), consequences for healthcare delivery (9%), near misses that would have led to harm if not for intervention (4%), no harm or consequences (3%), and complaints (2%). While most events involved device problems (93%), use problems (7%) were 4 times more likely to harm (relative risk 4.2; 95% CI 2.5-7). Problems with data input to ML devices were the top contributor to events (82%). DISCUSSION: Much of what is known about ML safety comes from case studies and the theoretical limitations of ML. We contribute a systematic analysis of ML safety problems captured as part of the FDA's routine post-market surveillance. Most problems involved devices and concerned the acquisition of data for processing by algorithms. However, problems with the use of devices were more likely to harm. CONCLUSIONS: Safety problems with ML devices involve more than algorithms, highlighting the need for a whole-of-system approach to safe implementation with a special focus on how users interact with devices. David Lyell, Ying Wang 0003, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Using automated methods to detect safety problems with health information technology: a scoping reviewabstractOBJECTIVE: To summarize the research literature evaluating automated methods for early detection of safety problems with health information technology (HIT). MATERIALS AND METHODS: We searched bibliographic databases including MEDLINE, ACM Digital, Embase, CINAHL Complete, PsycINFO, and Web of Science from January 2010 to June 2021 for studies evaluating the performance of automated methods to detect HIT problems. HIT problems were reviewed using an existing classification for safety concerns. Automated methods were categorized into rule-based, statistical, and machine learning methods, and their performance in detecting HIT problems was assessed. The review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta Analyses extension for Scoping Reviews statement. RESULTS: Of the 45 studies identified, the majority (n = 27, 60%) focused on detecting use errors involving electronic health records and order entry systems. Machine learning (n = 22) and statistical modeling (n = 17) were the most common methods. Unsupervised learning was used to detect use errors in laboratory test results, prescriptions, and patient records while supervised learning was used to detect technical errors arising from hardware or software issues. Statistical modeling was used to detect use errors, unauthorized access, and clinical decision support system malfunctions while rule-based methods primarily focused on use errors. CONCLUSIONS: A wide variety of rule-based, statistical, and machine learning methods have been applied to automate the detection of safety problems with HIT. Many opportunities remain to systematically study their application and effectiveness in real-world settings. Didi Surian, Ying Wang 0003, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Effects of machine learning-based clinical decision support systems on decision-making, care delivery, and patient outcomes: a scoping reviewabstractOBJECTIVE: This study aims to summarize the research literature evaluating machine learning (ML)-based clinical decision support (CDS) systems in healthcare settings. MATERIALS AND METHODS: We conducted a review in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta Analyses extension for Scoping Review). Four databases, including PubMed, Medline, Embase, and Scopus were searched for studies published from January 2016 to April 2021 evaluating the use of ML-based CDS in clinical settings. We extracted the study design, care setting, clinical task, CDS task, and ML method. The level of CDS autonomy was examined using a previously published 3-level classification based on the division of clinical tasks between the clinician and CDS; effects on decision-making, care delivery, and patient outcomes were summarized. RESULTS: Thirty-two studies evaluating the use of ML-based CDS in clinical settings were identified. All were undertaken in developed countries and largely in secondary and tertiary care settings. The most common clinical tasks supported by ML-based CDS were image recognition and interpretation (n = 12) and risk assessment (n = 9). The majority of studies examined assistive CDS (n = 23) which required clinicians to confirm or approve CDS recommendations for risk assessment in sepsis and for interpreting cancerous lesions in colonoscopy. Effects on decision-making, care delivery, and patient outcomes were mixed. CONCLUSION: ML-based CDS are being evaluated in many clinical areas. There remain many opportunities to apply and evaluate effects of ML-based CDS on decision-making, care delivery, and patient outcomes, particularly in resource-constrained settings. Anindya Pradipta Susanto, David Lyell, Bambang Widyantoro, Shlomo Berkovsky, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 5 |
| 2022 | More than Algorithms: an Analysis of ML Safety Events Reported to the FDA
David Lyell, Ying Wang 0003, Farah Magrabi |
AMIA | 3 |
| 2022 | What did you do to avoid the climate disaster? A call to arms for health informaticsabstractThe effects of human-induced climate change on our planet are already largely irreversible for many centuries1 and may remain so for at least 1000 years.2 If emissions continue to grow, their effects will trigger multiple critical tipping points and event cascades that will amplify climate effects in unpredicted ways.3 Just in 2022, we have seen flooding cover one-third of Pakistan, affecting 33 million people.4 India and Pakistan’s heatwave was the hottest yet on record.5 Recent years have seen historically extreme forest fires across North America, Europe, Asia, and Australia. Low-lying Pacific nations are slowly starting to disappear as sea levels rise, fed by melt waters from disappearing glaciers and sea ice, and coastal cities everywhere are at risk.6 The same emissions driving climate change are also affecting our health. Particulates in air pollution are likely responsible for around 300 000 lung cancer deaths globally,7 and the list of climate-induced health problems leading to poorer outcomes is depressingly long.8,9 Humanity is in trouble, and our way out is uncertain to say the least. Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 2 |
| 2022 | Digital health for climate change mitigation and response: a scoping reviewabstractOBJECTIVE: Climate change poses a major threat to the operation of global health systems, triggering large scale health events, and disrupting normal system operation. Digital health may have a role in the management of such challenges and in greenhouse gas emission reduction. This scoping review explores recent work on digital health responses and mitigation approaches to climate change. MATERIALS AND METHODS: We searched Medline up to February 11, 2022, using terms for digital health and climate change. Included articles were categorized into 3 application domains (mitigation, infectious disease, or environmental health risk management), and 6 technical tasks (data sensing, monitoring, electronic data capture, modeling, decision support, and communication). The review was PRISMA-ScR compliant. RESULTS: The 142 included publications reported a wide variety of research designs. Publication numbers have grown substantially in recent years, but few come from low- and middle-income countries. Digital health has the potential to reduce health system greenhouse gas emissions, for example by shifting to virtual services. It can assist in managing changing patterns of infectious diseases as well as environmental health events by timely detection, reducing exposure to risk factors, and facilitating the delivery of care to under-resourced areas. DISCUSSION: While digital health has real potential to help in managing climate change, research remains preliminary with little real-world evaluation. CONCLUSION: Significant acceleration in the quality and quantity of digital health climate change research is urgently needed, given the enormity of the global challenge. Hania Rahimi-Ardabili, Farah Magrabi, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 2 |
| 2021 | Automation in nursing decision support systems: A systematic review of effects on decision making, care delivery, and patient outcomesabstractOBJECTIVE: The study sought to summarize research literature on nursing decision support systems (DSSs ); understand which steps of the nursing care process (NCP) are supported by DSSs, and analyze effects of automated information processing on decision making, care delivery, and patient outcomes. MATERIALS AND METHODS: We conducted a systematic review in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) statement. PubMed, CINAHL, Cochrane, Embase, Scopus, and Web of Science were searched from January 2014 to April 2020 for studies focusing on DSSs used exclusively by nurses and their effects. Information about the stages of automation (information acquisition, information analysis, decision and action selection, and action implementation), NCP, and effects was assessed. RESULTS: Of 1019 articles retrieved, 28 met the inclusion criteria, each studying a unique DSS. Most DSSs were concerned with two NCP steps: assessment (82%) and intervention (86%). In terms of automation, all included DSSs automated information analysis and decision selection. Five DSSs automated information acquisition and only one automated action implementation. Effects on decision making, care delivery, and patient outcome were mixed. DSSs improved compliance with recommendations and reduced decision time, but impacts were not always sustainable. There were improvements in evidence-based practice, but impact on patient outcomes was mixed. CONCLUSIONS: Current nursing DSSs do not adequately support the NCP and have limited automation. There remain many opportunities to enhance automation, especially at the stage of information acquisition. Further research is needed to understand how automation within the NCP can improve nurses' decision making, care delivery, and patient outcomes. Saba Akbar, David Lyell, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | Evaluating the effects of automation on risk identification and nurses' decision making
Saba Akbar, David Lyell, Farah Magrabi |
AMIA | 3 |
| 2020 | Safety concerns with consumer-facing mobile health applications and their consequences: a scoping reviewabstractOBJECTIVE: To summarize the research literature about safety concerns with consumer-facing health apps and their consequences. MATERIALS AND METHODS: We searched bibliographic databases including PubMed, Web of Science, Scopus, and Cochrane libraries from January 2013 to May 2019 for articles about health apps. Descriptive information about safety concerns and consequences were extracted and classified into natural categories. The review was conducted in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) statement. RESULTS: Of the 74 studies identified, the majority were reviews of a single or a group of similar apps (n = 66, 89%), nearly half related to disease management (n = 34, 46%). A total of 80 safety concerns were identified, 67 related to the quality of information presented including incorrect or incomplete information, variation in content, and incorrect or inappropriate response to consumer needs. The remaining 13 related to app functionality including gaps in features, lack of validation for user input, delayed processing, failure to respond to health dangers, and faulty alarms. Of the 52 reports of actual or potential consequences, 5 had potential for patient harm. We also identified 66 reports about gaps in app development, including the lack of expert involvement, poor evidence base, and poor validation. CONCLUSIONS: Safety of apps is an emerging public health issue. The available evidence shows that apps pose clinical risks to consumers. Involvement of consumers, regulators, and healthcare professionals in development and testing can improve quality. Additionally, mandatory reporting of safety concerns is needed to improve outcomes. Saba Akbar, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 3 |
| 2020 | Can Unified Medical Language System-based semantic representation improve automated identification of patient safety incident reports by type and severity?abstractOBJECTIVE: The study sought to evaluate the feasibility of using Unified Medical Language System (UMLS) semantic features for automated identification of reports about patient safety incidents by type and severity. MATERIALS AND METHODS: Binary support vector machine (SVM) classifier ensembles were trained and validated using balanced datasets of critical incident report texts (n_type = 2860, n_severity = 1160) collected from a state-wide reporting system. Generalizability was evaluated on different and independent hospital-level reporting system. Concepts were extracted from report narratives using the UMLS Metathesaurus, and their relevance and frequency were used as semantic features. Performance was evaluated by F-score, Hamming loss, and exact match score and was compared with SVM ensembles using bag-of-words (BOW) features on 3 testing datasets (type/severity: n_benchmark = 286/116, n_original = 444/4837, n_independent =6000/5950). RESULTS: SVMs using semantic features met or outperformed those based on BOW features to identify 10 different incident types (F-score [semantics/BOW]: benchmark = 82.6%/69.4%; original = 77.9%/68.8%; independent = 78.0%/67.4%) and extreme-risk events (F-score [semantics/BOW]: benchmark = 87.3%/87.3%; original = 25.5%/19.8%; independent = 49.6%/52.7%). For incident type, the exact match score for semantic classifiers was consistently higher than BOW across all test datasets (exact match [semantics/BOW]: benchmark = 48.9%/39.9%; original = 57.9%/44.4%; independent = 59.5%/34.9%). DISCUSSION: BOW representations are not ideal for the automated identification of incident reports because they do not account for text semantics. UMLS semantic representations are likely to better capture information in report narratives, and thus may explain their superior performance. CONCLUSIONS: UMLS-based semantic classifiers were effective in identifying incidents by type and extreme-risk events, providing better generalizability than classifiers using BOW. Ying Wang 0003, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 3 |
| 2019 | Using convolutional neural networks to identify patient safety incident reports by type and severityabstractOBJECTIVE: To evaluate the feasibility of a convolutional neural network (CNN) with word embedding to identify the type and severity of patient safety incident reports. MATERIALS AND METHODS: A CNN with word embedding was applied to identify 10 incident types and 4 severity levels. Model training and validation used data sets (n_type = 2860, n_severity = 1160) collected from a statewide incident reporting system. Generalizability was evaluated using an independent hospital-level reporting system. CNN architectures were examined by varying layer size and hyperparameters. Performance was evaluated by F score, precision, recall, and compared to binary support vector machine (SVM) ensembles on 3 testing data sets (type/severity: n_benchmark = 286/116, n_original = 444/4837, n_independent = 6000/5950). RESULTS: A CNN with 6 layers was the most effective architecture, outperforming SVMs with better generalizability to identify incidents by type and severity. The CNN achieved high F scores (> 85%) across all test data sets when identifying common incident types including falls, medications, pressure injury, and aggression. When identifying common severity levels (medium/low), CNN outperformed SVMs, improving F scores by 11.9%-45.1% across all 3 test data sets. DISCUSSION: Automated identification of incident reports using machine learning is challenging because of a lack of large labelled training data sets and the unbalanced distribution of incident classes. The standard classification strategy is to build multiple binary classifiers and pool their predictions. CNNs can extract hierarchical features and assist in addressing class imbalance, which may explain their success in identifying incident report types. CONCLUSION: A CNN with word embedding was effective in identifying incidents by type and severity, providing better generalizability than SVMs. Ying Wang 0003, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 3 |
| 2018 | Safety concerns with consumer-facing mobile health applications and their consequences
Saba Akbar, Jessica A. Chen, Liliana Laranjo, Annie Y. S. Lau, Enrico W. Coiera, Farah Magrabi |
AMIA | 6 |
| 2018 | Technological Characteristics of Conversational Agents Used for Health-Related Purposes - A Systematic Review
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera |
AMIA | 9 |
| 2018 | Making Electronic Health Records Safer: Practical Strategies for Evaluation and Improvement
Allison B. McCoy, Dean F. Sittig, Adam Wright, Farah Magrabi |
AMIA | 4 |
| 2018 | Does health informatics have a replication crisis?abstractObjective: Many research fields, including psychology and basic medical sciences, struggle with poor reproducibility of reported studies. Biomedical and health informatics is unlikely to be immune to these challenges. This paper explores replication in informatics and the unique challenges the discipline faces. Methods: Narrative review of recent literature on research replication challenges. Results: While there is growing interest in re-analysis of existing data, experimental replication studies appear uncommon in informatics. Context effects are a particular challenge as they make ensuring replication fidelity difficult, and the same intervention will never quite reproduce the same result in different settings. Replication studies take many forms, trading-off testing validity of past findings against testing generalizability. Exact and partial replication designs emphasize testing validity while quasi and conceptual studies test generalizability of an underlying model or hypothesis with different methods or in a different setting. Conclusions: The cost of poor replication is a weakening in the quality of published research and the evidence-based foundation of health informatics. The benefits of replication include increased rigor in research, and the development of evaluation methods that distinguish the impact of context and the nonreproducibility of research. Taking replication seriously is essential if biomedical and health informatics is to be an evidence-based discipline. Enrico W. Coiera, Elske Ammenwerth, Andrew Georgiou, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 4 |
| 2018 | Conversational agents in healthcare: a systematic reviewabstractObjective: Our objective was to review the characteristics, current applications, and evaluation measures of conversational agents with unconstrained natural language input capabilities used for health-related purposes. Methods: We searched PubMed, Embase, CINAHL, PsycInfo, and ACM Digital using a predefined search strategy. Studies were included if they focused on consumers or healthcare professionals; involved a conversational agent using any unconstrained natural language input; and reported evaluation measures resulting from user interaction with the system. Studies were screened by independent reviewers and Cohen's kappa measured inter-coder agreement. Results: The database search retrieved 1513 citations; 17 articles (14 different conversational agents) met the inclusion criteria. Dialogue management strategies were mostly finite-state and frame-based (6 and 7 conversational agents, respectively); agent-based strategies were present in one type of system. Two studies were randomized controlled trials (RCTs), 1 was cross-sectional, and the remaining were quasi-experimental. Half of the conversational agents supported consumers with health tasks such as self-care. The only RCT evaluating the efficacy of a conversational agent found a significant effect in reducing depression symptoms (effect size d = 0.44, p = .04). Patient safety was rarely evaluated in the included studies. Conclusions: The use of conversational agents with unconstrained natural language input capabilities for health-related purposes is an emerging field of research, where the few published studies were mainly quasi-experimental, and rarely evaluated efficacy or safety. Future studies would benefit from more robust experimental designs and standardized reporting. Protocol Registration: The protocol for this systematic review is registered at PROSPERO with the number CRD42017065917. Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 9 |
| 2017 | Engineering technology resilience through informatics safety scienceabstractWith every year that passes, our relationship to information technology becomes more complex, and our dependence deeper. Technology is our great ally, promising greater efficiency and productivity. It also promises greater safety for our patients. However, this relationship with technology can sometimes be a brittle one. We can quickly cross a safety gap from a comfortable place where everything works well, to one where the limits of technology introduce new risks. Whether it is through a computer network failure, applying a system software patch, or a user accidentally clicking on the wrong patient name, it is surprisingly easy to move from safe to unsafe. As the footprint of technology across our health services has grown, so to by extrapolation, has the associated risk of technology harms to patients.1 It is the potential abruptness of this transition to increased risk of harm, this lack of graceful degradation in performance, and the silence accompanying degradation, that remain unsolved challenges to the effective use of information technology in healthcare. Enrico W. Coiera, Farah Magrabi, Jan L. Talmon |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Efficiency and safety of speech recognition for documentation in the electronic health recordabstractOBJECTIVE: To compare the efficiency and safety of using speech recognition (SR) assisted clinical documentation within an electronic health record (EHR) system with use of keyboard and mouse (KBM). METHODS: Thirty-five emergency department clinicians undertook randomly allocated clinical documentation tasks using KBM or SR on a commercial EHR system. Tasks were simple or complex, and with or without interruption. Outcome measures included task completion times and observed errors. Errors were classed by their potential for patient harm. Error causes were classified as due to IT system/system integration, user interaction, comprehension, or as typographical. User-related errors could be by either omission or commission. RESULTS: Mean task completion times were 18.11% slower overall when using SR compared to KBM (P = .001), 16.95% slower for simple tasks (P = .050), and 18.40% slower for complex tasks (P = .009). Increased errors were observed with use of SR (KBM 32, SR 138) for both simple (KBM 9, SR 75; P < 0.001) and complex (KBM 23, SR 63; P < 0.001) tasks. Interruptions did not significantly affect task completion times or error rates for either modality. DISCUSSION: For clinical documentation, SR was slower and increased the risk of documentation errors, including errors with the potential to cause clinical harm compared to KBM. Some of the observed increase in errors may be due to suboptimal SR to EHR integration and workflow. CONCLUSION: Use of SR to drive interactive clinical documentation in the EHR requires careful evaluation. Current generation implementations may require significant development before they are safe and effective. Improving system integration and workflow, as well as SR accuracy and user-focused error correction strategies, may improve SR performance. Tobias Hodgson, Farah Magrabi, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Problems with health information technology and their effects on care delivery and patient outcomes: a systematic reviewabstractObjective: To systematically review studies reporting problems with information technology (IT) in health care and their effects on care delivery and patient outcomes. Materials and methods: We searched bibliographic databases including Scopus, PubMed, and Science Citation Index Expanded from January 2004 to December 2015 for studies reporting problems with IT and their effects. A framework called the information value chain, which connects technology use to final outcome, was used to assess how IT problems affect user interaction, information receipt, decision-making, care processes, and patient outcomes. The review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) statement. Results: Of the 34 studies identified, the majority ( n = 14, 41%) were analyses of incidents reported from 6 countries. There were 7 descriptive studies, 9 ethnographic studies, and 4 case reports. The types of IT problems were similar to those described in earlier classifications of safety problems associated with health IT. The frequency, scale, and severity of IT problems were not adequately captured within these studies. Use errors and poor user interfaces interfered with the receipt of information and led to errors of commission when making decisions. Clinical errors involving medications were well characterized. Issues with system functionality, including poor user interfaces and fragmented displays, delayed care delivery. Issues with system access, system configuration, and software updates also delayed care. In 18 studies (53%), IT problems were linked to patient harm and death. Near-miss events were reported in 10 studies (29%). Discussion and conclusion: The research evidence describing problems with health IT remains largely qualitative, and many opportunities remain to systematically study and quantify risks and benefits with regard to patient safety. The information value chain, when used in conjunction with existing classifications for health IT safety problems, can enhance measurement and should facilitate identification of the most significant risks to patient safety. Mi Ok Kim, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 3 |
| 2016 | Measuring the effects of computer downtime on hospital pathology processes
Ying Wang 0003, Enrico W. Coiera, Blanca Gallego, Óscar Pérez, Mei-Sing Ong, Guy Tsafnat, David Roffe, Graham Jones, Farah Magrabi |
J. Biomed. Informatics | 9 |
| 2015 | Health information technology and large-scale adverse events
Farah Magrabi, Dean F. Sittig, Jean M. Scott, Peter M. Kilbridge |
AMIA | 1 |
| 2013 | Health IT: Assessment of Safe and Effective Use - Measuring the HIT Hazard Function
Blackford Middleton, Christian Nøhr, James Walker, Farah Magrabi |
AMIA | 4 |
| 2013 | Using statistical text classification to identify health information technology incidentsabstractOBJECTIVE: To examine the feasibility of using statistical text classification to automatically identify health information technology (HIT) incidents in the USA Food and Drug Administration (FDA) Manufacturer and User Facility Device Experience (MAUDE) database. DESIGN: We used a subset of 570 272 incidents including 1534 HIT incidents reported to MAUDE between 1 January 2008 and 1 July 2010. Text classifiers using regularized logistic regression were evaluated with both 'balanced' (50% HIT) and 'stratified' (0.297% HIT) datasets for training, validation, and testing. Dataset preparation, feature extraction, feature selection, cross-validation, classification, performance evaluation, and error analysis were performed iteratively to further improve the classifiers. Feature-selection techniques such as removing short words and stop words, stemming, lemmatization, and principal component analysis were examined. MEASUREMENTS: κ statistic, F1 score, precision and recall. RESULTS: Classification performance was similar on both the stratified (0.954 F1 score) and balanced (0.995 F1 score) datasets. Stemming was the most effective technique, reducing the feature set size to 79% while maintaining comparable performance. Training with balanced datasets improved recall (0.989) but reduced precision (0.165). CONCLUSIONS: Statistical text classification appears to be a feasible method for identifying HIT reports within large databases of incidents. Automated identification should enable more HIT problems to be detected, analyzed, and addressed in a timely manner. Semi-supervised learning may be necessary when applying machine learning to big data analysis of patient safety incidents and requires further investigation. Kevin E. K. Chai, Stephen Anthony, Enrico W. Coiera, Farah Magrabi |
J. Am. Medical Informatics Assoc. | 4 |
| 2013 | Syndromic surveillance for health information system failures: a feasibility studyabstractOBJECTIVE: To explore the applicability of a syndromic surveillance method to the early detection of health information technology (HIT) system failures. METHODS: A syndromic surveillance system was developed to monitor a laboratory information system at a tertiary hospital. Four indices were monitored: (1) total laboratory records being created; (2) total records with missing results; (3) average serum potassium results; and (4) total duplicated tests on a patient. The goal was to detect HIT system failures causing: data loss at the record level; data loss at the field level; erroneous data; and unintended duplication of data. Time-series models of the indices were constructed, and statistical process control charts were used to detect unexpected behaviors. The ability of the models to detect HIT system failures was evaluated using simulated failures, each lasting for 24 h, with error rates ranging from 1% to 35%. RESULTS: In detecting data loss at the record level, the model achieved a sensitivity of 0.26 when the simulated error rate was 1%, while maintaining a specificity of 0.98. Detection performance improved with increasing error rates, achieving a perfect sensitivity when the error rate was 35%. In the detection of missing results, erroneous serum potassium results and unintended repetition of tests, perfect sensitivity was attained when the error rate was as small as 5%. Decreasing the error rate to 1% resulted in a drop in sensitivity to 0.65-0.85. CONCLUSIONS: Syndromic surveillance methods can potentially be applied to monitor HIT systems, to facilitate the early detection of failures. Mei-Sing Ong, Farah Magrabi, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | National and cross-border safety initiatives for health information technology
Farah Magrabi, Dean F. Sittig, Maureen Baker, Jan L. Talmon, Enrico W. Coiera |
AMIA | 1 |
| 2012 | Impact of a web-based personally controlled health management system on influenza vaccination and health services utilization rates: a randomized controlled trialabstractOBJECTIVE: To assess the impact of a web-based personally controlled health management system (PCHMS) on the uptake of seasonal influenza vaccine and primary care service utilization among university students and staff. MATERIALS AND METHODS: A PCHMS called Healthy.me was developed and evaluated in a 2010 CONSORT-compliant two-group (6-month waitlist vs PCHMS) parallel randomized controlled trial (RCT) (allocation ratio 1:1). The PCHMS integrated an untethered personal health record with consumer care pathways, social forums, and messaging links with a health service provider. RESULTS: 742 university students and staff met inclusion criteria and were randomized to a 6-month waitlist (n=372) or the PCHMS (n=370). Amongst the 470 participants eligible for primary analysis, PCHMS users were 6.7% (95% CI: 1.46 to 12.30) more likely than the waitlist to receive an influenza vaccine (waitlist: 4.9% (12/246, 95% CI 2.8 to 8.3) vs PCHMS: 11.6% (26/224, 95% CI 8.0 to 16.5); χ(2)=7.1, p=0.008). PCHMS participants were also 11.6% (95% CI 3.6 to 19.5) more likely to visit the health service provider (waitlist: 17.9% (44/246, 95% CI 13.6 to 23.2) vs PCHMS: 29.5% (66/224, 95% CI: 23.9 to 35.7); χ(2)=8.8, p=0.003). A dose-response effect was detected, where greater use of the PCHMS was associated with higher rates of vaccination (p=0.001) and health service provider visits (p=0.003). DISCUSSION: PCHMS can significantly increase consumer participation in preventive health activities, such as influenza vaccination. CONCLUSIONS: Integrating a PCHMS into routine health service delivery systems appears to be an effective mechanism for enhancing consumer engagement in preventive health measures. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12610000386033. http://www.anzctr.org.au/trial_view.aspx?id=335463. Annie Y. S. Lau, Vitali Sintchenko, Jacinta Crimmins, Farah Magrabi, Blanca Gallego, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 4 |
| 2012 | A systematic review of the psychological literature on interruption and its patient safety implicationsabstractOBJECTIVE: To understand the complex effects of interruption in healthcare. MATERIALS AND METHODS: As interruptions have been well studied in other domains, the authors undertook a systematic review of experimental studies in psychology and human-computer interaction to identify the task types and variables influencing interruption effects. RESULTS: 63 studies were identified from 812 articles retrieved by systematic searches. On the basis of interruption profiles for generic tasks, it was found that clinical tasks can be distinguished into three broad types: procedural, problem-solving, and decision-making. Twelve experimental variables that influence interruption effects were identified. Of these, six are the most important, based on the number of studies and because of their centrality to interruption effects, including working memory load, interruption position, similarity, modality, handling strategies, and practice effect. The variables are explained by three main theoretical frameworks: the activation-based goal memory model, prospective memory, and multiple resource theory. DISCUSSION: This review provides a useful starting point for a more comprehensive examination of interruptions potentially leading to an improved understanding about the impact of this phenomenon on patient safety and task efficiency. The authors provide some recommendations to counter interruption effects. CONCLUSION: The effects of interruption are the outcome of a complex set of variables and should not be considered as uniformly predictable or bad. The task types, variables, and theories should help us better to identify which clinical tasks and contexts are most susceptible and assist in the design of information systems and processes that are resilient to interruption. Simon Y. W. Li, Farah Magrabi, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 2 |
| 2012 | Using FDA reports to inform a classification for health information technology safety problemsabstractOBJECTIVE: To expand an emerging classification for problems with health information technology (HIT) using reports submitted to the US Food and Drug Administration Manufacturer and User Facility Device Experience (MAUDE) database. DESIGN: HIT events submitted to MAUDE were retrieved using a standardized search strategy. Using an emerging classification with 32 categories of HIT problems, a subset of relevant events were iteratively analyzed to identify new categories. Two coders then independently classified the remaining events into one or more categories. Free-text descriptions were analyzed to identify the consequences of events. MEASUREMENTS: Descriptive statistics by number of reported problems per category and by consequence; inter-rater reliability analysis using the κ statistic for the major categories and consequences. RESULTS: A search of 899 768 reports from January 2008 to July 2010 yielded 1100 reports about HIT. After removing duplicate and unrelated reports, 678 reports describing 436 events remained. The authors identified four new categories to describe problems with software functionality, system configuration, interface with devices, and network configuration; the authors' classification with 32 categories of HIT problems was expanded by the addition of these four categories. Examination of the 436 events revealed 712 problems, 96% were machine-related, and 4% were problems at the human-computer interface. Almost half (46%) of the events related to hazardous circumstances. Of the 46 events (11%) associated with patient harm, four deaths were linked to HIT problems (0.9% of 436 events). CONCLUSIONS: Only 0.1% of the MAUDE reports searched were related to HIT. Nevertheless, Food and Drug Administration reports did prove to be a useful new source of information about the nature of software problems and their safety implications with potential to inform strategies for safe design and implementation. Farah Magrabi, Mei-Sing Ong, William Runciman, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Automated identification of extreme-risk events in clinical incident reportsabstractOBJECTIVES: To explore the feasibility of using statistical text classification to automatically detect extreme-risk events in clinical incident reports. METHODS: Statistical text classifiers based on Naïve Bayes and Support Vector Machine (SVM) algorithms were trained and tested on clinical incident reports to automatically detect extreme-risk events, defined by incidents that satisfy the criteria of Severity Assessment Code (SAC) level 1. For this purpose, incident reports submitted to the Advanced Incident Management System by public hospitals from one Australian region were used. The classifiers were evaluated on two datasets: (1) a set of reports with diverse incident types (n=120); (2) a set of reports associated with patient misidentification (n=166). Results were assessed using accuracy, precision, recall, F-measure, and area under the curve (AUC) of receiver operating characteristic curves. RESULTS: The classifiers performed well on both datasets. In the multi-type dataset, SVM with a linear kernel performed best, identifying 85.8% of SAC level 1 incidents (precision=0.88, recall=0.83, F-measure=0.86, AUC=0.92). In the patient misidentification dataset, 96.4% of SAC level 1 incidents were detected when SVM with linear, polynomial or radial-basis function kernel was used (precision=0.99, recall=0.94, F-measure=0.96, AUC=0.98). Naïve Bayes showed reasonable performance, detecting 80.8% of SAC level 1 incidents in the multi-type dataset and 89.8% of SAC level 1 patient misidentification incidents. Overall, higher prediction accuracy was attained on the specialized dataset, compared with the multi-type dataset. CONCLUSION: Text classification techniques can be applied effectively to automate the detection of extreme-risk events in clinical incident reports. Mei-Sing Ong, Farah Magrabi, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 2 |
| 2011 | A simulation framework for mapping risks in clinical processes: the case of in-patient transfersabstractOBJECTIVE: To model how individual violations in routine clinical processes cumulatively contribute to the risk of adverse events in hospital using an agent-based simulation framework. DESIGN: An agent-based simulation was designed to model the cascade of common violations that contribute to the risk of adverse events in routine clinical processes. Clinicians and the information systems that support them were represented as a group of interacting agents using data from direct observations. The model was calibrated using data from 101 patient transfers observed in a hospital and results were validated for one of two scenarios (a misidentification scenario and an infection control scenario). Repeated simulations using the calibrated model were undertaken to create a distribution of possible process outcomes. The likelihood of end-of-chain risk is the main outcome measure, reported for each of the two scenarios. RESULTS: The simulations demonstrate end-of-chain risks of 8% and 24% for the misidentification and infection control scenarios, respectively. Over 95% of the simulations in both scenarios are unique, indicating that the in-patient transfer process diverges from prescribed work practices in a variety of ways. CONCLUSIONS: The simulation allowed us to model the risk of adverse events in a clinical process, by generating the variety of possible work subject to violations, a novel prospective risk analysis method. The in-patient transfer process has a high proportion of unique trajectories, implying that risk mitigation may benefit from focusing on reducing complexity rather than augmenting the process with further rule-based protocols. Adam G. Dunn, Mei-Sing Ong, Johanna I. Westbrook, Farah Magrabi, Enrico W. Coiera, Wayne Wobcke |
J. Am. Medical Informatics Assoc. | 4 |
| 2010 | Errors and electronic prescribing: a controlled laboratory study to examine task complexity and interruption effectsabstractOBJECTIVE: To examine the effect of interruptions and task complexity on error rates when prescribing with computerized provider order entry (CPOE) systems, and to categorize the types of prescribing errors. DESIGN: Two within-subject factors: task complexity (complex vs simple) and interruption (interruption vs no interruption). Thirty-two hospital doctors used a CPOE system in a computer laboratory to complete four prescribing tasks, half of which were interrupted using a counterbalanced design. MEASUREMENTS: Types of prescribing errors, error rate, resumption lag, and task completion time. RESULTS: Errors in creating and updating electronic medication charts that were measured included failure to enter allergy information; selection of incorrect medication, dose, route, formulation, or frequency of administration from lists and drop-down menus presented by the CPOE system; incorrect entry or omission in entering administration times, start date, and free-text qualifiers; and omissions in prescribing and ceasing medications. When errors occurred, the error rates across the four prescribing tasks ranged from 0.5% (1 incorrect medication selected out of 192 chances for selecting a medication or error opportunities) to 16% (5 failures to enter allergy information out of 32 error opportunities). Any impact of interruptions on prescribing error rates and task completion times was not detected in our experiment. However, complex tasks took significantly longer to complete (F(1, 27)=137.9; p<0.001) and when execution was interrupted they required almost three times longer to resume compared to simple tasks (resumption lag complex=9.6 seconds, SD=5.6; resumption lag simple=3.4 seconds, SD=1.7; t(28)=6.186; p<0.001). CONCLUSION: Most electronic prescribing errors found in this study could be described as slips in using the CPOE system to create and update electronic medication charts. Cues available within the user interface may have aided resumption of interrupted tasks making CPOE systems robust to some interruption effects. Further experiments are required to rule out any effect interruption might have on CPOE error rates. Farah Magrabi, Simon Y. W. Li, Richard O. Day, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 1 |
| 2010 | An analysis of computer-related patient safety incidents to inform the development of a classificationabstractOBJECTIVE: To analyze patient safety incidents associated with computer use to develop the basis for a classification of problems reported by health professionals. DESIGN: Incidents submitted to a voluntary incident reporting database across one Australian state were retrieved and a subset (25%) was analyzed to identify 'natural categories' for classification. Two coders independently classified the remaining incidents into one or more categories. Free text descriptions were analyzed to identify contributing factors. Where available medical specialty, time of day and consequences were examined. MEASUREMENTS: Descriptive statistics; inter-rater reliability. RESULTS: A search of 42,616 incidents from 2003 to 2005 yielded 123 computer related incidents. After removing duplicate and unrelated incidents, 99 incidents describing 117 problems remained. A classification with 32 types of computer use problems was developed. Problems were grouped into information input (31%), transfer (20%), output (20%) and general technical (24%). Overall, 55% of problems were machine related and 45% were attributed to human-computer interaction. Delays in initiating and completing clinical tasks were a major consequence of machine related problems (70%) whereas rework was a major consequence of human-computer interaction problems (78%). While 38% (n=26) of the incidents were reported to have a noticeable consequence but no harm, 34% (n=23) had no noticeable consequence. CONCLUSION: Only 0.2% of all incidents reported were computer related. Further work is required to expand our classification using incident reports and other sources of information about healthcare IT problems. Evidence based user interface design must focus on the safe entry and retrieval of clinical information and support users in detecting and correcting errors and malfunctions. Farah Magrabi, Mei-Sing Ong, William Runciman, Enrico W. Coiera |
J. Am. Medical Informatics Assoc. | 1 |