Zina M. Ibrahim

dblp:55/6241 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0001-6203-2727ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 CSAI: Conditional Self-Attention Imputation for Healthcare Time-Series
abstract
We introduce the Conditional Self-Attention Imputation (CSAI) model, a novel recurrent neural network architecture designed to address imputation challenges in multivariate time series derived from hospital electronic health records (EHRs). CSAI introduces key novelties specific to EHR data: a) attention-based hidden state initialisation to capture both long- and short-range temporal dependencies, b) domain-informed temporal decay to mimic clinical recording patterns, and c) a non-uniform masking strategy that models non-random missingness. Comprehensive evaluation across four EHR benchmark datasets demonstrates CSAI's effectiveness compared to state-of-the-art architectures in data restoration and downstream tasks. CSAI is integrated into PyPOTS, an open-source Python toolbox for partially observed time series. This work significantly advances the state of neural network imputation applied to EHRs by more closely aligning algorithmic imputation with clinical realities.
Linglong Qian, Joseph Arul Raj, Hugh Logan Ellis, Yuezhou Zhang 0001, Tao Wang 0036, Richard J. B. Dobson, Zina M. Ibrahim
IEEE J. Biomed. Health Informatics8
2025 How Deep is Your Guess? A Fresh Perspective on Deep Learning for Medical Time-Series Imputation
abstract
We present a comprehensive analysis of deep learning approaches for Electronic Health Record (EHR) time-series imputation, examining how the interplay between architectural and framework design decisions gives rise to higher-level properties of a given deep imputer model and distinct biases towards complex data characteristics. Our investigation reveals the varying capabilities of deep imputers in capturing complex spatio-temporal dependencies within EHRs, and that the effectiveness of the model depends on how its combined biases align with the characteristics of the medical time series. Our experimental evaluation challenges common assumptions about model complexity, demonstrating that larger models do not necessarily improve performance. Rather, carefully designed architectures can better capture the complex patterns inherent in clinical data. The study highlights the need for imputation approaches that prioritise clinically meaningful data reconstruction over statistical accuracy. Our experiments further reveal up to 20% in variations of imputation performance based on preprocessing and implementation choices, emphasising the need for standardised benchmarking methodologies. Finally, we identify critical gaps between current deep imputation methods and medical requirements, highlighting the importance of integrating clinical insights to achieve more reliable imputation approaches for healthcare applications.
Linglong Qian, Hugh Logan Ellis, Tao Wang 0036, Jun Wang 0121, Robin Mitra, Richard J. B. Dobson, Zina M. Ibrahim
IEEE J. Biomed. Health Informatics7
2023 Discharge summary hospital course summarisation of in patient Electronic Health Record text with clinical concept guided deep pre-trained Transformer models
Thomas Searle, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson
J. Biomed. Informatics2
2022 A Knowledge Distillation Ensemble Framework for Predicting Short- and Long-Term Hospitalization Outcomes From Electronic Health Records Data
abstract
The ability to perform accurate prognosis is crucial for proactive clinical decision making, informed resource management and personalised care. Existing outcome prediction models suffer from a low recall of infrequent positive outcomes. We present a highly-scalable and robust machine learning framework to automatically predict adversity represented by mortality and ICU admission and readmission from time-series of vital signs and laboratory results obtained within the first 24 hours of hospital admission. The stacked ensemble platform comprises two components: a) an unsupervised LSTM Autoencoder that learns an optimal representation of the time-series, using it to differentiate the less frequent patterns which conclude with an adverse event from the majority patterns that do not, and b) a gradient boosting model, which relies on the constructed representation to refine prediction by incorporating static features. The model is used to assess a patient's risk of adversity and provides visual justifications of its prediction. Results of three case studies show that the model outperforms existing platforms in ICU and general ward settings, achieving average Precision-Recall Areas Under the Curve (PR-AUCs) of 0.891 (95% CI: 0.878-0.939) for mortality and 0.908 (95% CI: 0.870-0.935) in predicting ICU admission and readmission.
Zina M. Ibrahim, Daniel Bean, Thomas Searle, Linglong Qian, Honghan Wu, Anthony Shek, Zeljko Kraljevic, James Galloway, Sam Norton, James T. Teo, Richard J. B. Dobson
IEEE J. Biomed. Health Informatics1
2021 Multi-domain clinical natural language processing with MedCAT: The Medical Concept Annotation Toolkit
Zeljko Kraljevic, Thomas Searle, Anthony Shek, Lukasz Roguski, Kawsar Noor, Daniel Bean, Aurelie Mascio, Leilei Zhu, Amos Folarin, Angus Roberts, Rebecca Bendayan, Mark P. Richardson, Robert Stewart 0002, Anoop D. Shah, Wai Keong Wong, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson
Artif. Intell. Medicine16
2021 Ensemble learning for poor prognosis predictions: A case study on SARS-CoV-2
abstract
OBJECTIVE: Risk prediction models are widely used to inform evidence-based clinical decision making. However, few models developed from single cohorts can perform consistently well at population level where diverse prognoses exist (such as the SARS-CoV-2 [severe acute respiratory syndrome coronavirus 2] pandemic). This study aims at tackling this challenge by synergizing prediction models from the literature using ensemble learning. MATERIALS AND METHODS: In this study, we selected and reimplemented 7 prediction models for COVID-19 (coronavirus disease 2019) that were derived from diverse cohorts and used different implementation techniques. A novel ensemble learning framework was proposed to synergize them for realizing personalized predictions for individual patients. Four diverse international cohorts (2 from the United Kingdom and 2 from China; N = 5394) were used to validate all 8 models on discrimination, calibration, and clinical usefulness. RESULTS: Results showed that individual prediction models could perform well on some cohorts while poorly on others. Conversely, the ensemble model achieved the best performances consistently on all metrics quantifying discrimination, calibration, and clinical usefulness. Performance disparities were observed in cohorts from the 2 countries: all models achieved better performances on the China cohorts. DISCUSSION: When individual models were learned from complementary cohorts, the synergized model had the potential to achieve better performances than any individual model. Results indicate that blood parameters and physiological measurements might have better predictive powers when collected early, which remains to be confirmed by further studies. CONCLUSIONS: Combining a diverse set of individual prediction models, the ensemble method can synergize a robust and well-performing model by choosing the most competent ones for individual patients.
Honghan Wu, Andreas Karwath, Zina M. Ibrahim, Kevin Dhaliwal, Daniel Bean, Victor Roth Cardoso, Kezhi Li, James T. Teo, Amitava Banerjee, Fang Gao-Smith, Tony Whitehouse, Tonny Veenith, Georgios V. Gkoutos, Richard J. B. Dobson, Bruce Guthrie
J. Am. Medical Informatics Assoc.4
2021 Estimating redundancy in clinical text
Thomas Searle, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson
J. Biomed. Informatics2
2020 Modeling Rare Interactions in Time Series Data Through Qualitative Change: Application to Outcome Prediction in Intensive Care Units
abstract
Many areas of research are characterised by the deluge of large-scale highly-dimensional time-series data. However, using the data available for prediction and decision making is hampered by the current lag in our ability to uncover and quantify true interactions that explain the outcomes. We are interested in areas such as intensive care medicine, which are characterised by i) continuous monitoring of multivariate variables and non-uniform sampling of data streams, ii) the outcomes are generally governed by interactions between a small set of rare events, iii) these interactions are not necessarily definable by specific values (or value ranges) of a given group of variables, but rather, by the deviations of these values from the normal state recorded over time, iv) the need to explain the predictions made by the model. Here, while numerous data mining models have been formulated for outcome prediction, they are unable to explain their predictions. We present a model for uncovering interactions with the highest likelihood of generating the outcomes seen from highly-dimensional time series data. Interactions among variables are represented by a relational graph structure, which relies on qualitative abstractions to overcome non-uniform sampling and to capture the semantics of the interactions corresponding to the changes and deviations from normality of variables of interest over time. Using the assumption that similar templates of small interactions are responsible for the outcomes (as prevalent in the medical domains), we reformulate the discovery task to retrieve the most-likely templates from the data. Experiments on sepsis prediction using real Intensive Care Unit (ICU) data demonstrates that the discovered interaction templates are semantically meaningful within the domain, and using them as features in a prediction task produces a superior performance than when using the raw values of the predictors.
Zina M. Ibrahim, Honghan Wu, Richard J. B. Dobson
ECAI1
2020 Comparing Natural Language Processing Techniques for Alzheimer's Dementia Prediction in Spontaneous Speech
abstract
Alzheimer's Dementia (AD) is an incurable, debilitating, and progressive neurodegenerative condition that affects cognitive function. Early diagnosis is important as therapeutics can delay progression and give those diagnosed vital time. Developing models that analyse spontaneous speech could eventually provide an efficient diagnostic modality for earlier diagnosis of AD. The Alzheimer's Dementia Recognition through Spontaneous Speech task offers acoustically pre-processed and balanced datasets for the classification and prediction of AD and associated phenotypes through the modelling of spontaneous speech. We exclusively analyse the supplied textual transcripts of the spontaneous speech dataset, building and comparing performance across numerous models for the classification of AD vs controls and the prediction of Mental Mini State Exam scores. We rigorously train and evaluate Support Vector Machines (SVMs), Gradient Boosting Decision Trees (GBDT), and Conditional Random Fields (CRFs) alongside deep learning Transformer based models. We find our top performing models to be a simple Term Frequency-Inverse Document Frequency (TF-IDF) vectoriser as input into a SVM model and a pre-trained Transformer based model `DistilBERT' when used as an embedding layer into simple linear models. We demonstrate test set scores of 0.81-0.82 across classification metrics and a RMSE of 4.58.
Thomas Searle, Zina M. Ibrahim, Richard J. B. Dobson
INTERSPEECH2
2020 On classifying sepsis heterogeneity in the ICU: insight using machine learning
abstract
OBJECTIVES: Current machine learning models aiming to predict sepsis from electronic health records (EHR) do not account 20 for the heterogeneity of the condition despite its emerging importance in prognosis and treatment. This work demonstrates the added value of stratifying the types of organ dysfunction observed in patients who develop sepsis in the intensive care unit (ICU) in improving the ability to recognize patients at risk of sepsis from their EHR data. MATERIALS AND METHODS: Using an ICU dataset of 13 728 records, we identify clinically significant sepsis subpopulations with distinct organ dysfunction patterns. We perform classification experiments with random forest, gradient boost trees, and support vector machines, using the identified subpopulations to distinguish patients who develop sepsis in the ICU from those who do not. RESULTS: The classification results show that features selected using sepsis subpopulations as background knowledge yield a superior performance in distinguishing septic from non-septic patients regardless of the classification model used. The improved performance is especially pronounced in specificity, which is a current bottleneck in sepsis prediction machine learning models. CONCLUSION: Our findings can steer machine learning efforts toward more personalized models for complex conditions including sepsis.
Zina M. Ibrahim, Honghan Wu, Ahmed Hamoud, Lukas Stappen, Richard J. B. Dobson, Andrea Agarossi
J. Am. Medical Informatics Assoc.1
2018 SemEHR: A general-purpose semantic search system to surface semantic data from clinical notes for tailored care, trial recruitment, and clinical research
abstract
Objective: Unlocking the data contained within both structured and unstructured components of electronic health records (EHRs) has the potential to provide a step change in data available for secondary research use, generation of actionable medical insights, hospital management, and trial recruitment. To achieve this, we implemented SemEHR, an open source semantic search and analytics tool for EHRs. Methods: SemEHR implements a generic information extraction (IE) and retrieval infrastructure by identifying contextualized mentions of a wide range of biomedical concepts within EHRs. Natural language processing annotations are further assembled at the patient level and extended with EHR-specific knowledge to generate a timeline for each patient. The semantic data are serviced via ontology-based search and analytics interfaces. Results: SemEHR has been deployed at a number of UK hospitals, including the Clinical Record Interactive Search, an anonymized replica of the EHR of the UK South London and Maudsley National Health Service Foundation Trust, one of Europe's largest providers of mental health services. In 2 Clinical Record Interactive Search-based studies, SemEHR achieved 93% (hepatitis C) and 99% (HIV) F-measure results in identifying true positive patients. At King's College Hospital in London, as part of the CogStack program (github.com/cogstack), SemEHR is being used to recruit patients into the UK Department of Health 100 000 Genomes Project (genomicsengland.co.uk). The validation study suggests that the tool can validate previously recruited cases and is very fast at searching phenotypes; time for recruitment criteria checking was reduced from days to minutes. Validated on open intensive care EHR data, Medical Information Mart for Intensive Care III, the vital signs extracted by SemEHR can achieve around 97% accuracy. Conclusion: Results from the multiple case studies demonstrate SemEHR's efficiency: weeks or months of work can be done within hours or minutes in some cases. SemEHR provides a more comprehensive view of patients, bringing in more and unexpected insight compared to study-oriented bespoke IE systems. SemEHR is open source, available at https://github.com/CogStack/SemEHR.
Honghan Wu, Giulia Toti, Katherine Morley, Zina M. Ibrahim, Amos Folarin, Richard G. Jackson, Ismail Emre Kartoglu, Asha Agrawal, Clive Stringer, Darren Gale, Genevieve Gorrell, Angus Roberts, Matthew T. M. Broadbent, Robert Stewart 0002, Richard J. B. Dobson
J. Am. Medical Informatics Assoc.4
2015 The relative vertex clustering value - a new criterion for the fast discovery of functional modules in protein interaction networks
abstract
BACKGROUND: Cellular processes are known to be modular and are realized by groups of proteins implicated in common biological functions. Such groups of proteins are called functional modules, and many community detection methods have been devised for their discovery from protein interaction networks (PINs) data. In current agglomerative clustering approaches, vertices with just a very few neighbors are often classified as separate clusters, which does not make sense biologically. Also, a major limitation of agglomerative techniques is that their computational efficiency do not scale well to large PINs. Finally, PIN data obtained from large scale experiments generally contain many false positives, and this makes it hard for agglomerative clustering methods to find the correct clusters, since they are known to be sensitive to noisy data. RESULTS: We propose a local similarity premetric, the relative vertex clustering value, as a new criterion allowing to decide when a node can be added to a given node's cluster and which addresses the above three issues. Based on this criterion, we introduce a novel and very fast agglomerative clustering technique, FAC-PIN, for discovering functional modules and protein complexes from a PIN data. CONCLUSIONS: Our proposed FAC-PIN algorithm is applied to nine PIN data from eight different species including the yeast PIN, and the identified functional modules are validated using Gene Ontology (GO) annotations from DAVID Bioinformatics Resources. Identified protein complexes are also validated using experimentally verified complexes. Computational results show that FAC-PIN can discover functional modules or protein complexes from PINs more accurately and more efficiently than HC-PIN and CNM, the current state-of-the-art approaches for clustering PINs in an agglomerative manner.
Zina M. Ibrahim, Alioune Ngom
BMC Bioinform.1
2013 Detecting epistasis in the presence of linkage disequilibrium: A focused comparison
abstract
We present results from a comparison of three epistasis-detection tools using large-scale simulated genetic data: SNPHarvester, SNPRuler and Ambience. The tools were chosen based on their merits to be representative of the state of the art of epistasis detection. We design and conduct experiments to test the performance of the methods in detecting interacting loci or their proxies in linkage disequilibrium (LD) tagged regions, in datasets containing simulated 2,3 and 4-way epistatic interactions. The results show that SNPHarvester is the fastest while Ambience is the most robust. Moreover, SNPRuler provides the best power, specially with higher-level interactions, but cannot scale-up to larger datasets.
Zina M. Ibrahim, Stephen Newhouse, Richard J. B. Dobson
CIBCB1
2011 Using Qualitative Probability in Reverse-Engineering Gene Regulatory Networks
abstract
This paper demonstrates the use of qualitative probabilistic networks (QPNs) to aid Dynamic Bayesian Networks (DBNs) in the process of learning the structure of gene regulatory networks from microarray gene expression data. We present a study which shows that QPNs define monotonic relations that are capable of identifying regulatory interactions in a manner that is less susceptible to the many sources of uncertainty that surround gene expression data. Moreover, we construct a model that maps the regulatory interactions of genetic networks to QPN constructs and show its capability in providing a set of candidate regulators for target genes, which is subsequently used to establish a prior structure that the DBN learning algorithm can use and which 1) distinguishes spurious correlations from true regulations, 2) enables the discovery of sets of coregulators of target genes, and 3) results in a more efficient construction of gene regulatory networks. The model is compared to the existing literature using the known gene regulatory interactions of Drosophila Melanogaster.
Zina M. Ibrahim, Alioune Ngom, Ahmed Y. Tawfik
IEEE ACM Trans. Comput. Biol. Bioinform.1
2010 A dynamic qualitative probabilistic network approach for extracting gene regulatory network motifs
abstract
This paper extends our work to using qualitative probability to model the naturally-occurring motifs of gene regulatory networks. Having showed in [16] that the qualitative relations defining QPN graphs exhibit a direct mapping to the naturally-occurring network motifs embedded in Gene Regulatory Networks, this work is concerned with generalizing QPN constructs to create a high-level framework from which any regulatory network motif can be derived. Experimental results using time-series data of the Saccharomyces Cerevisiae show the effectiveness of our approach in providing a more accurate description of the regulatory motifs in the Saccharomyces Cerevisiae gene regulatory network compared to our previous definitions.
Zina M. Ibrahim, Alioune Ngom, Ahmed Y. Tawfik
BIBM1
2009 Qualitative Motif Detection in Gene Regulatory Networks
abstract
This paper motivates the use of qualitative probabilistic networks (QPNs) in conjunction with or in lieu of Bayesian Networks (BNs) for reconstructing gene regulatory networks from microarray expression data. QPNs are qualitative abstractions of Bayesian Networks that replace the conditional probability tables associated with BNs by qualitative influences, which use signs to encode how the values of variables change. We demonstrate that the qualitative influences defined by QPNs exhibit a natural mapping to naturally-occurring patterns of connections, termed network motifs, embedded in Gene Regulatory Networks and present a model that maps QPN constructs to such motifs. The contribution of this paper is that of discovering motifs by mapping their time-series experimental data to QPN influences and using the discovered motifs to aid the process of reconstructing the corresponding gene regulatory network via Dynamic Bayesian Networks (DBNs). The general aim is to compile a model that uses qualitative equivalents of Dynamic Bayesian Networks to explore gene expression networks and their regulatory mechanisms. Although this aim remains under development, the results we have obtained shows success for the discovery of regulatory motifs in Saccharomyces Cerevisiae and their effectiveness in improving the results obtained in terms of reconstruction using DBNs.
Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom
BIBM1
2009 Surprise-Based Qualitative Probabilistic Networks
Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom
ECSQARU1
2007 A Qualitative Hidden Markov Model for Spatio-temporal Reasoning
Zina M. Ibrahim, Ahmed Y. Tawfik, Alioune Ngom
ECSQARU1