Blanca Gallego

dblp:66/7236 · also Blanca Gallego Luxan · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-3704-7975ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Using natural language processing to extract information from clinical text in electronic medical records for populating clinical registries: a systematic review
abstract
OBJECTIVE: Clinical registries advance healthcare by tracking patient outcomes and intervention safety. Manually extracting information from clinical text for registries is labor- and resource-intensive and often inaccurate. Therefore, this systematic review aims to evaluate the use and effectiveness of natural language processing (NLP) methods in extracting information from clinical text for populating clinical registries. MATERIALS AND METHODS: PubMed, Embase, Scopus, Web of Science, and ACM Digital Library were systematically searched. Studies were included if they used NLP techniques to populate clinical registries. The extracted data included details of the registry, the clinical text, the registry data elements extracted, the NLP methods used, and how their performance was evaluated. RESULTS: Fifteen articles were included in the review. Since 2020, the use of NLP methods for extracting information to populate clinical registries has been increasing steadily. Initially, rule-based NLP methods dominated the field, but machine learning-based approaches have gradually gained popularity. However, only one of the included studies employed generative large language models (LLMs). The diversity of clinical text and extracted data elements posed challenges to the generalizability of the NLP methods. CONCLUSION: To date, the application of NLP methods to clinical text for populating clinical registries has been limited in both the number of published studies and the scope of implementation. The NLP methods used thus far face significant challenges in effectively managing the complexity and diversity of clinical text and data elements. Moreover, the performance of the NLP methods varied significantly. This review underscores the need for a robust and adaptable NLP framework. Generative LLMs may provide direction for future research, but their use must account for challenges such as accuracy, cost, privacy, and limited supporting evidence.
Leibo Liu, Victoria Blake, Matthew Barman, Blanca Gallego, Timothy Churches, Georgina Kennedy, Sze-Yuan Ooi, Geoffrey Delaney, Louisa Jorm
J. Am. Medical Informatics Assoc.4
2025 Attention-based synthetic data generation for calibration-enhanced survival analysis: A case study for chronic kidney disease using electronic health records
abstract
OBJECTIVES: Access to real-world healthcare data is constrained by privacy regulations and data imbalances, hindering the development of fair and reliable clinical prediction models. Synthetic data offers a potential solution, yet existing methods often fail to maintain calibration or enable subgroup-specific augmentation. This study introduces Masked Clinical Modelling (MCM), an attention-based synthetic data generation framework designed to enhance survival model calibration in both global and stratified analyses. METHODS: MCM uses masked feature reconstruction to learn feature dependencies without explicitly training on survival objectives. It supports both standalone dataset synthesis and conditional data augmentation, enabling the generation of targeted synthetic subcohorts without retraining. Evaluated on a chronic kidney disease (CKD) electronic health record (EHR) dataset, MCM was benchmarked against eight baseline methods, including variational autoencoders, GANs, SMOTE variants, and a recent risk-aware distillation model. Model performance was assessed via calibration loss, Cox model consistency, and Kaplan-Meier fidelity. RESULTS: MCM-generated data closely replicated statistical properties of the real dataset, pre- served hazard ratios, and matched time-to-event curves with high fidelity. Cox models trained on MCM-augmented data demonstrated improved calibration, reducing overall calibration loss by 15% and subgroup meta-calibration loss by 9% compared to unaugmented data. These improvements held across multiple high-risk subgroups including those with diabetes, renal dys- function, and advanced age. Unlike competing methods, MCM achieved this without retraining or outcome-specific tuning. CONCLUSIONS: MCM offers a practical and flexible framework for generating synthetic survival data that improves risk model calibration. By supporting both reproducible dataset synthesis and conditional subgroup augmentation, MCM bridges privacy-preserving data access with calibration-aware learning. This work highlights the role of synthetic data not just as a privacy tool, but as a vehicle for improving equity and reliability in clinical modelling.
Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm
J. Biomed. Informatics2
2023 Continuous time recurrent neural networks: Overview and benchmarking at forecasting blood glucose in the intensive care unit
abstract
OBJECTIVE: Blood glucose measurements in the intensive care unit (ICU) are typically made at irregular intervals. This presents a challenge in choice of forecasting model. This article gives an overview of continuous time autoregressive recurrent neural networks (CTRNNs) and evaluates how they compare to autoregressive gradient boosted trees (GBT) in forecasting blood glucose in the ICU. METHODS: Continuous time autoregressive recurrent neural networks (CTRNNs) are a deep learning model that account for irregular observations through incorporating continuous evolution of the hidden states between observations. This is achieved using a neural ordinary differential equation (ODE) or neural flow layer. In this manuscript, we give an overview of these models, including the varying architectures that have been proposed to account for issues such as ongoing medical interventions. Further, we demonstrate the application of these models to probabilistic forecasting of blood glucose in a critical care setting using electronic medical record and simulated data and compare with GBT and linear models. RESULTS: The experiments confirm that addition of a neural ODE or neural flow layer generally improves the performance of autoregressive recurrent neural networks in the irregular measurement setting. However, several CTRNN architecture are outperformed by a GBT model (Catboost), with only a long short-term memory (LSTM) and neural ODE based architecture (ODE-LSTM) achieving comparable performance on probabilistic forecasting metrics such as the continuous ranked probability score (ODE-LSTM: 0.118 ± 0.001; Catboost: 0.118 ± 0.001), ignorance score (0.152 ± 0.008; 0.149 ± 0.002) and interval score (175 ± 1; 176 ± 1). CONCLUSION: The application of deep learning methods for forecasting in situations with irregularly measured time series such as blood glucose shows promise. However, appropriate benchmarking by methods such as GBT approaches (plus feature transformation) are key in highlighting whether novel methodologies are truly state of the art in tabular data settings.
Oisin Fitzgerald, Óscar Pérez, Blanca Gallego, Alejandro Metke-Jimenez, Lachlan Rudd, Louisa Jorm
J. Biomed. Informatics3
2022 Causal inference for observational longitudinal studies using deep survival models
abstract
OBJECTIVE: Causal inference for observational longitudinal studies often requires the accurate estimation of treatment effects on time-to-event outcomes in the presence of time-dependent patient history and time-dependent covariates. MATERIALS AND METHODS: To tackle this longitudinal treatment effect estimation problem, we have developed a time-variant causal survival (TCS) model that uses the potential outcomes framework with an ensemble of recurrent subnetworks to estimate the difference in survival probabilities and its confidence interval over time as a function of time-dependent covariates and treatments. RESULTS: Using simulated survival datasets, the TCS model showed good causal effect estimation performance across scenarios of varying sample dimensions, event rates, confounding and overlapping. However, increasing the sample size was not effective in alleviating the adverse impact of a high level of confounding. In a large clinical cohort study, TCS identified the expected conditional average treatment effect and detected individual treatment effect heterogeneity over time. TCS provides an efficient way to estimate and update individualized treatment effects over time, in order to improve clinical decisions. DISCUSSION: The use of a propensity score layer and potential outcome subnetworks helps correcting for selection bias. However, the proposed model is limited in its ability to correct the bias from unmeasured confounding, and more extensive testing of TCS under extreme scenarios such as low overlapping and the presence of unmeasured confounders is desired and left for future work. CONCLUSION: TCS fills the gap in causal inference using deep learning techniques in survival analysis. It considers time-varying confounders and treatment options. Its treatment effect estimation can be easily compared with the conventional literature, which uses relative measures of treatment effect. We expect TCS will be particularly useful for identifying and quantifying treatment effect heterogeneity over time under the ever complex observational health care environment.
Blanca Gallego
J. Biomed. Informatics2
2021 Incorporating real-world evidence into the development of patient blood glucose prediction algorithms for the ICU
abstract
OBJECTIVE: Glycemic control is an important component of critical care. We present a data-driven method for predicting intensive care unit (ICU) patient response to glycemic control protocols while accounting for patient heterogeneity and variations in care. MATERIALS AND METHODS: Using electronic medical records (EMRs) of 18 961 ICU admissions from the MIMIC-III dataset, including 318 574 blood glucose measurements, we train and validate a gradient boosted tree machine learning (ML) algorithm to forecast patient blood glucose and a 95% prediction interval at 2-hour intervals. The model uses as inputs irregular multivariate time series data relating to recent in-patient medical history and glycemic control, including previous blood glucose, nutrition, and insulin dosing. RESULTS: Our forecasting model using routinely collected EMRs achieves performance comparable to previous models developed in planned research studies using continuous blood glucose monitoring. Model error, expressed as mean absolute percentage error is 16.5%-16.8%, with Clarke error grid analysis demonstrating that 97% of predictions would be clinically acceptable. The 95% prediction intervals achieve near intended coverage at 93%-94%. DISCUSSION: ML algorithms built on observational data sources, such as EMRs, present a promising approach for personalization and automation of glycemic control in critical care. Future research may benefit from applying a combination of methodologies and data sources to develop robust methodologies that account for the variations seen in ICU patients and difficultly in detecting the extremes of observed blood glucose values. CONCLUSION: We demonstrate that EMRs can be used to train ML algorithms that may be suitable for incorporation into ICU decision support systems.
Oisin Fitzgerald, Óscar Pérez, Blanca Gallego, Manoj K. Saxena, Lachlan Rudd, Alejandro Metke-Jimenez, Louisa Jorm
J. Am. Medical Informatics Assoc.3
2020 Targeted estimation of heterogeneous treatment effect in observational survival analysis
Blanca Gallego
J. Biomed. Informatics2
2018 Technological Characteristics of Conversational Agents Used for Health-Related Purposes - A Systematic Review
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera
AMIA8
2018 Conversational agents in healthcare: a systematic review
abstract
Objective: Our objective was to review the characteristics, current applications, and evaluation measures of conversational agents with unconstrained natural language input capabilities used for health-related purposes. Methods: We searched PubMed, Embase, CINAHL, PsycInfo, and ACM Digital using a predefined search strategy. Studies were included if they focused on consumers or healthcare professionals; involved a conversational agent using any unconstrained natural language input; and reported evaluation measures resulting from user interaction with the system. Studies were screened by independent reviewers and Cohen's kappa measured inter-coder agreement. Results: The database search retrieved 1513 citations; 17 articles (14 different conversational agents) met the inclusion criteria. Dialogue management strategies were mostly finite-state and frame-based (6 and 7 conversational agents, respectively); agent-based strategies were present in one type of system. Two studies were randomized controlled trials (RCTs), 1 was cross-sectional, and the remaining were quasi-experimental. Half of the conversational agents supported consumers with health tasks such as self-care. The only RCT evaluating the efficacy of a conversational agent found a significant effect in reducing depression symptoms (effect size d = 0.44, p = .04). Patient safety was rarely evaluated in the included studies. Conclusions: The use of conversational agents with unconstrained natural language input capabilities for health-related purposes is an emerging field of research, where the few published studies were mainly quasi-experimental, and rarely evaluated efficacy or safety. Future studies would benefit from more robust experimental designs and standardized reporting. Protocol Registration: The protocol for this systematic review is registered at PROSPERO with the number CRD42017065917.
Liliana Laranjo, Adam G. Dunn, Huong Ly Tong, Ahmet Baki Kocaballi, Jessica A. Chen, Rabia Bashir, Didi Surian, Blanca Gallego, Farah Magrabi, Annie Y. S. Lau, Enrico W. Coiera
J. Am. Medical Informatics Assoc.8
2016 Real-time prediction of mortality, readmission, and length of stay using electronic health record data
abstract
OBJECTIVE: To develop a predictive model for real-time predictions of length of stay, mortality, and readmission for hospitalized patients using electronic health records (EHRs). MATERIALS AND METHODS: A Bayesian Network model was built to estimate the probability of a hospitalized patient being "at home," in the hospital, or dead for each of the next 7 days. The network utilizes patient-specific administrative and laboratory data and is updated each time a new pathology test result becomes available. Electronic health records from 32 634 patients admitted to a Sydney metropolitan hospital via the emergency department from July 2008 through December 2011 were used. The model was tested on 2011 data and trained on the data of earlier years. RESULTS: The model achieved an average daily accuracy of 80% and area under the receiving operating characteristic curve (AUROC) of 0.82. The model's predictive ability was highest within 24 hours from prediction (AUROC = 0.83) and decreased slightly with time. Death was the most predictable outcome with a daily average accuracy of 93% and AUROC of 0.84. DISCUSSION: We developed the first non-disease-specific model that simultaneously predicts remaining days of hospitalization, death, and readmission as part of the same outcome. By providing a future daily probability for each outcome class, we enable the visualization of future patient trajectories. Among these, it is possible to identify trajectories indicating expected discharge, expected continuing hospitalization, expected death, and possible readmission. CONCLUSIONS: Bayesian Networks can model EHRs to provide real-time forecasts for patient outcomes, which provide richer information than traditional independent point predictions of length of stay, death, or readmission, and can thus better support decision making.
Xiongcai Cai, Óscar Pérez, Enrico W. Coiera, Fernando Martín-Sánchez, Richard O. Day, David Roffe, Blanca Gallego
J. Am. Medical Informatics Assoc.7
2016 Measuring the effects of computer downtime on hospital pathology processes
Ying Wang 0003, Enrico W. Coiera, Blanca Gallego, Óscar Pérez, Mei-Sing Ong, Guy Tsafnat, David Roffe, Graham Jones, Farah Magrabi
J. Biomed. Informatics3
2012 Impact of a web-based personally controlled health management system on influenza vaccination and health services utilization rates: a randomized controlled trial
abstract
OBJECTIVE: To assess the impact of a web-based personally controlled health management system (PCHMS) on the uptake of seasonal influenza vaccine and primary care service utilization among university students and staff. MATERIALS AND METHODS: A PCHMS called Healthy.me was developed and evaluated in a 2010 CONSORT-compliant two-group (6-month waitlist vs PCHMS) parallel randomized controlled trial (RCT) (allocation ratio 1:1). The PCHMS integrated an untethered personal health record with consumer care pathways, social forums, and messaging links with a health service provider. RESULTS: 742 university students and staff met inclusion criteria and were randomized to a 6-month waitlist (n=372) or the PCHMS (n=370). Amongst the 470 participants eligible for primary analysis, PCHMS users were 6.7% (95% CI: 1.46 to 12.30) more likely than the waitlist to receive an influenza vaccine (waitlist: 4.9% (12/246, 95% CI 2.8 to 8.3) vs PCHMS: 11.6% (26/224, 95% CI 8.0 to 16.5); χ(2)=7.1, p=0.008). PCHMS participants were also 11.6% (95% CI 3.6 to 19.5) more likely to visit the health service provider (waitlist: 17.9% (44/246, 95% CI 13.6 to 23.2) vs PCHMS: 29.5% (66/224, 95% CI: 23.9 to 35.7); χ(2)=8.8, p=0.003). A dose-response effect was detected, where greater use of the PCHMS was associated with higher rates of vaccination (p=0.001) and health service provider visits (p=0.003). DISCUSSION: PCHMS can significantly increase consumer participation in preventive health activities, such as influenza vaccination. CONCLUSIONS: Integrating a PCHMS into routine health service delivery systems appears to be an effective mechanism for enhancing consumer engagement in preventive health measures. TRIAL REGISTRATION: Australian New Zealand Clinical Trials Registry ACTRN12610000386033. http://www.anzctr.org.au/trial_view.aspx?id=335463.
Annie Y. S. Lau, Vitali Sintchenko, Jacinta Crimmins, Farah Magrabi, Blanca Gallego, Enrico W. Coiera
J. Am. Medical Informatics Assoc.5
2009 Towards bioinformatics assisted infectious disease control
abstract
BACKGROUND: This paper proposes a novel framework for bioinformatics assisted biosurveillance and early warning to address the inefficiencies in traditional surveillance as well as the need for more timely and comprehensive infection monitoring and control. It leverages on breakthroughs in rapid, high-throughput molecular profiling of microorganisms and text mining. RESULTS: This framework combines the genetic and geographic data of a pathogen to reconstruct its history and to identify the migration routes through which the strains spread regionally and internationally. A pilot study of Salmonella typhimurium genotype clustering and temporospatial outbreak analysis demonstrated better discrimination power than traditional phage typing. Half of the outbreaks were detected in the first half of their duration. CONCLUSION: The microbial profiling and biosurveillance focused text mining tools can enable integrated infectious disease outbreak detection and response environments based upon bioinformatics knowledge models and measured by outcomes including the accuracy and timeliness of outbreak detection.
Vitali Sintchenko, Blanca Gallego, Grace Chung, Enrico W. Coiera
BMC Bioinform.2
2009 Biosurveillance of emerging biothreats using scalable genotype clustering
Blanca Gallego, Vitali Sintchenko, Qinning Wang, Lester Hiley, Gwendolyn L. Gilbert, Enrico W. Coiera
J. Biomed. Informatics1