EDBT 2026 Demo / reviewers in the wild / expert
Louisa Jorm
dblp:241/7267 · also Louisa R. Jorm
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0003-0390-661XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Using natural language processing to extract information from clinical text in electronic medical records for populating clinical registries: a systematic reviewabstractOBJECTIVE: Clinical registries advance healthcare by tracking patient outcomes and intervention safety. Manually extracting information from clinical text for registries is labor- and resource-intensive and often inaccurate. Therefore, this systematic review aims to evaluate the use and effectiveness of natural language processing (NLP) methods in extracting information from clinical text for populating clinical registries. MATERIALS AND METHODS: PubMed, Embase, Scopus, Web of Science, and ACM Digital Library were systematically searched. Studies were included if they used NLP techniques to populate clinical registries. The extracted data included details of the registry, the clinical text, the registry data elements extracted, the NLP methods used, and how their performance was evaluated. RESULTS: Fifteen articles were included in the review. Since 2020, the use of NLP methods for extracting information to populate clinical registries has been increasing steadily. Initially, rule-based NLP methods dominated the field, but machine learning-based approaches have gradually gained popularity. However, only one of the included studies employed generative large language models (LLMs). The diversity of clinical text and extracted data elements posed challenges to the generalizability of the NLP methods. CONCLUSION: To date, the application of NLP methods to clinical text for populating clinical registries has been limited in both the number of published studies and the scope of implementation. The NLP methods used thus far face significant challenges in effectively managing the complexity and diversity of clinical text and data elements. Moreover, the performance of the NLP methods varied significantly. This review underscores the need for a robust and adaptable NLP framework. Generative LLMs may provide direction for future research, but their use must account for challenges such as accuracy, cost, privacy, and limited supporting evidence. Leibo Liu, Victoria Blake, Matthew Barman, Blanca Gallego, Timothy Churches, Georgina Kennedy, Sze-Yuan Ooi, Geoffrey Delaney, Louisa Jorm |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | Attention-based synthetic data generation for calibration-enhanced survival analysis: A case study for chronic kidney disease using electronic health recordsabstractOBJECTIVES: Access to real-world healthcare data is constrained by privacy regulations and data imbalances, hindering the development of fair and reliable clinical prediction models. Synthetic data offers a potential solution, yet existing methods often fail to maintain calibration or enable subgroup-specific augmentation. This study introduces Masked Clinical Modelling (MCM), an attention-based synthetic data generation framework designed to enhance survival model calibration in both global and stratified analyses. METHODS: MCM uses masked feature reconstruction to learn feature dependencies without explicitly training on survival objectives. It supports both standalone dataset synthesis and conditional data augmentation, enabling the generation of targeted synthetic subcohorts without retraining. Evaluated on a chronic kidney disease (CKD) electronic health record (EHR) dataset, MCM was benchmarked against eight baseline methods, including variational autoencoders, GANs, SMOTE variants, and a recent risk-aware distillation model. Model performance was assessed via calibration loss, Cox model consistency, and Kaplan-Meier fidelity. RESULTS: MCM-generated data closely replicated statistical properties of the real dataset, pre- served hazard ratios, and matched time-to-event curves with high fidelity. Cox models trained on MCM-augmented data demonstrated improved calibration, reducing overall calibration loss by 15% and subgroup meta-calibration loss by 9% compared to unaugmented data. These improvements held across multiple high-risk subgroups including those with diabetes, renal dys- function, and advanced age. Unlike competing methods, MCM achieved this without retraining or outcome-specific tuning. CONCLUSIONS: MCM offers a practical and flexible framework for generating synthetic survival data that improves risk model calibration. By supporting both reproducible dataset synthesis and conditional subgroup augmentation, MCM bridges privacy-preserving data access with calibration-aware learning. This work highlights the role of synthetic data not just as a privacy tool, but as a vehicle for improving equity and reliability in clinical modelling. Nicholas I-Hsien Kuo, Blanca Gallego, Louisa Jorm |
J. Biomed. Informatics | 3 |
| 2023 | Automated ICD coding using extreme multi-label long text transformer-based modelsabstractEncouraged by the success of pretrained Transformer models in many natural language processing tasks, their use for International Classification of Diseases (ICD) coding tasks is now actively being explored. In this study, we investigated two existing Transformer-based models (PLM-ICD and XR-Transformer) and proposed a novel Transformer-based model (XR-LAT), aiming to address the extreme label set and long text classification challenges that are posed by automated ICD coding tasks. The Transformer-based model PLM-ICD, which currently holds the state-of-the-art (SOTA) performance on the ICD coding benchmark datasets MIMIC-III and MIMIC-II, was selected as our baseline model for further optimisation on both datasets. In addition, we extended the capabilities of the leading model in the general extreme multi-label text classification domain, XR-Transformer, to support longer sequences and trained it on both datasets. Moreover, we proposed a novel model, XR-LAT, which was also trained on both datasets. XR-LAT is a recursively trained model chain on a predefined hierarchical code tree with label-wise attention, knowledge transferring and dynamic negative sampling mechanisms. Our optimised PLM-ICD models, which were trained with longer total and chunk sequence lengths, significantly outperformed the current SOTA PLM-ICD models, and achieved the highest micro-F1 scores of 60.8 % and 50.9 % on MIMIC-III and MIMIC-II, respectively. The XR-Transformer model, although SOTA in the general domain, did not perform well across all metrics. The best XR-LAT based models obtained results that were competitive with the current SOTA PLM-ICD models, including improving the macro-AUC by 2.1 % and 5.1 % on MIMIC-III and MIMIC-II, respectively. Our optimised PLM-ICD models are the new SOTA models for automated ICD coding on both datasets, while our novel XR-LAT models perform competitively with the previous SOTA PLM-ICD models. Leibo Liu, Óscar Pérez, Anthony N. Nguyen, Vicki Bennett, Louisa Jorm |
Artif. Intell. Medicine | 5 |
| 2023 | Continuous time recurrent neural networks: Overview and benchmarking at forecasting blood glucose in the intensive care unitabstractOBJECTIVE: Blood glucose measurements in the intensive care unit (ICU) are typically made at irregular intervals. This presents a challenge in choice of forecasting model. This article gives an overview of continuous time autoregressive recurrent neural networks (CTRNNs) and evaluates how they compare to autoregressive gradient boosted trees (GBT) in forecasting blood glucose in the ICU. METHODS: Continuous time autoregressive recurrent neural networks (CTRNNs) are a deep learning model that account for irregular observations through incorporating continuous evolution of the hidden states between observations. This is achieved using a neural ordinary differential equation (ODE) or neural flow layer. In this manuscript, we give an overview of these models, including the varying architectures that have been proposed to account for issues such as ongoing medical interventions. Further, we demonstrate the application of these models to probabilistic forecasting of blood glucose in a critical care setting using electronic medical record and simulated data and compare with GBT and linear models. RESULTS: The experiments confirm that addition of a neural ODE or neural flow layer generally improves the performance of autoregressive recurrent neural networks in the irregular measurement setting. However, several CTRNN architecture are outperformed by a GBT model (Catboost), with only a long short-term memory (LSTM) and neural ODE based architecture (ODE-LSTM) achieving comparable performance on probabilistic forecasting metrics such as the continuous ranked probability score (ODE-LSTM: 0.118 ± 0.001; Catboost: 0.118 ± 0.001), ignorance score (0.152 ± 0.008; 0.149 ± 0.002) and interval score (175 ± 1; 176 ± 1). CONCLUSION: The application of deep learning methods for forecasting in situations with irregularly measured time series such as blood glucose shows promise. However, appropriate benchmarking by methods such as GBT approaches (plus feature transformation) are key in highlighting whether novel methodologies are truly state of the art in tabular data settings. Oisin Fitzgerald, Óscar Pérez, Blanca Gallego, Alejandro Metke-Jimenez, Lachlan Rudd, Louisa Jorm |
J. Biomed. Informatics | 6 |
| 2023 | Generating synthetic clinical data that capture class imbalanced distributions with generative adversarial networks: Example using antiretroviral therapy for HIVabstractOBJECTIVE: Clinical data's confidential nature often limits the development of machine learning models in healthcare. Generative adversarial networks (GANs) can synthesise realistic datasets, but suffer from mode collapse, resulting in low diversity and bias towards majority demographics and common clinical practices. This work proposes an extension to the classic GAN framework that includes a variational autoencoder (VAE) and an external memory mechanism to overcome these limitations and generate synthetic data accurately describing imbalanced class distributions commonly found in clinical variables. METHODS: The proposed method generated a synthetic dataset related to antiretroviral therapy for human immunodeficiency virus (ART for HIV). We evaluated it based on five metrics: (1) accurately representing imbalanced class distribution; (2) the realism of the individual variables; (3) the realism among variables; (4) patient disclosure risk; and (5) the utility of the generated dataset for developing downstream machine learning models. RESULTS: The proposed method overcomes the issue of mode collapse and generates a synthetic dataset that accurately describes imbalanced class distributions commonly found in clinical variables. The generated data has a patient disclosure risk of 0.095%, lower than the 9% threshold stated by Health Canada and the European Medicines Agency, making it suitable for distribution to the research community with high security. The generated data also has high utility, indicating the potential of the proposed method to enable the development of downstream machine learning algorithms for healthcare applications using synthetic data. CONCLUSION: Our proposed extension to the classic GAN framework, which includes a VAE and an external memory mechanism, represents a promising approach towards generating synthetic data that accurately describe imbalanced class distributions commonly found in clinical variables. This method overcomes the limitations of GANs and creates more realistic datasets with higher patient cohort diversity, facilitating the development of downstream machine learning algorithms for healthcare applications. Nicholas I-Hsien Kuo, Federico Garcia, Anders Sönnerborg, Michael Böhm, Rolf Kaiser, Maurizio Zazzi, Mark N. Polizzotto, Louisa Jorm, Sebastiano Barbieri |
J. Biomed. Informatics | 8 |
| 2022 | Hierarchical label-wise attention transformer model for explainable ICD codingabstractInternational Classification of Diseases (ICD) coding plays an important role in systematically classifying morbidity and mortality data. In this study, we propose a hierarchical label-wise attention Transformer model (HiLAT) for the explainable prediction of ICD codes from clinical documents. HiLAT firstly fine-tunes a pretrained Transformer model to represent the tokens of clinical documents. We subsequently employ a two-level hierarchical label-wise attention mechanism that creates label-specific document representations. These representations are in turn used by a feed-forward neural network to predict whether a specific ICD code is assigned to the input clinical document of interest. We evaluate HiLAT using hospital discharge summaries and their corresponding ICD-9 codes from the MIMIC-III database. To investigate the performance of different types of Transformer models, we develop ClinicalplusXLNet, which conducts continual pretraining from XLNet-Base using all the MIMIC-III clinical notes. The experiment results show that the F1 scores of the HiLAT + ClinicalplusXLNet outperform the previous state-of-the-art models for the top-50 most frequent ICD-9 codes from MIMIC-III. Visualisations of attention weights present a potential explainability tool for checking the face validity of ICD code predictions. Leibo Liu, Óscar Pérez, Anthony N. Nguyen, Vicki Bennett, Louisa Jorm |
J. Biomed. Informatics | 5 |
| 2022 | De-identifying Australian hospital discharge summaries: An end-to-end framework using ensemble of deep learning modelsabstractElectronic Medical Records (EMRs) contain clinical narrative text that is of great potential value to medical researchers. However, this information is mixed with Personally Identifiable Information (PII) that presents risks to patient and clinician confidentiality. This paper presents an end-to-end de-identification framework to automatically remove PII from Australian hospital discharge summaries. Our corpus included 600 hospital discharge summaries which were extracted from the EMRs of two principal referral hospitals in Sydney, Australia. Our end-to-end de-identification framework consists of three components: (1) Annotation: labelling of PII in the 600 hospital discharge summaries using five pre-defined categories: person, address, date of birth, individual identification number, phone/fax number; (2) Modelling: training six named entity recognition (NER) deep learning base-models on balanced and imbalanced datasets; and evaluating ensembles that combine all six base-models, the three base-models with the best F1 scores and the three base-models with the best recall scores respectively, using token-level majority voting and stacking methods; and (3) De-identification: removing PII from the hospital discharge summaries. Our results showed that the ensemble model combined using the stacking Support Vector Machine (SVM) method on the three base-models with the best F1 scores achieved excellent results with a F1 score of 99.16% on the test set of our corpus. We also evaluated the robustness of our modelling component on the 2014 i2b2 de-identification dataset. Our ensemble model, which uses the token-level majority voting method on all six base-models, achieved the highest F1 score of 96.24% at strict entity matching and the highest F1 score of 98.64% at binary token-level matching compared to two state-of-the-art methods. The end-to-end framework provides a robust solution to de-identifying clinical narrative corpuses safely. It can easily be applied to any kind of clinical narrative documents. Leibo Liu, Óscar Pérez, Anthony N. Nguyen, Vicki Bennett, Louisa Jorm |
J. Biomed. Informatics | 5 |
| 2021 | Incorporating real-world evidence into the development of patient blood glucose prediction algorithms for the ICUabstractOBJECTIVE: Glycemic control is an important component of critical care. We present a data-driven method for predicting intensive care unit (ICU) patient response to glycemic control protocols while accounting for patient heterogeneity and variations in care. MATERIALS AND METHODS: Using electronic medical records (EMRs) of 18 961 ICU admissions from the MIMIC-III dataset, including 318 574 blood glucose measurements, we train and validate a gradient boosted tree machine learning (ML) algorithm to forecast patient blood glucose and a 95% prediction interval at 2-hour intervals. The model uses as inputs irregular multivariate time series data relating to recent in-patient medical history and glycemic control, including previous blood glucose, nutrition, and insulin dosing. RESULTS: Our forecasting model using routinely collected EMRs achieves performance comparable to previous models developed in planned research studies using continuous blood glucose monitoring. Model error, expressed as mean absolute percentage error is 16.5%-16.8%, with Clarke error grid analysis demonstrating that 97% of predictions would be clinically acceptable. The 95% prediction intervals achieve near intended coverage at 93%-94%. DISCUSSION: ML algorithms built on observational data sources, such as EMRs, present a promising approach for personalization and automation of glycemic control in critical care. Future research may benefit from applying a combination of methodologies and data sources to develop robust methodologies that account for the variations seen in ICU patients and difficultly in detecting the extremes of observed blood glucose values. CONCLUSION: We demonstrate that EMRs can be used to train ML algorithms that may be suitable for incorporation into ICU decision support systems. Oisin Fitzgerald, Óscar Pérez, Blanca Gallego, Manoj K. Saxena, Lachlan Rudd, Alejandro Metke-Jimenez, Louisa Jorm |
J. Am. Medical Informatics Assoc. | 7 |