VLDB 2026 Research / reviewers in the wild / expert
Richard J. B. Dobson
dblp:51/5120
· DBLP profile ↗
32ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-4224-9245ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 9 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VIEWER: an extensible visual analytics framework for enhancing mental healthcareabstractOBJECTIVE: A proof-of-concept study aimed at designing and implementing Visual & Interactive Engagement With Electronic Records (VIEWER), a versatile toolkit for visual analytics of clinical data, and systematically evaluating its effectiveness across various clinical applications while gathering feedback for iterative improvements. MATERIALS AND METHODS: VIEWER is an open-source and extensible toolkit that employs natural language processing and interactive visualization techniques to facilitate the rapid design, development, and deployment of clinical information retrieval, analysis, and visualization at the point of care. Through an iterative and collaborative participatory design approach, VIEWER was designed and implemented in one of the United Kingdom's largest National Health Services mental health Trusts, where its clinical utility and effectiveness were assessed using both quantitative and qualitative methods. RESULTS: VIEWER provides interactive, problem-focused, and comprehensive views of longitudinal patient data (n = 409 870) from a combination of structured clinical data and unstructured clinical notes. Despite a relatively short adoption period and users' initial unfamiliarity, VIEWER significantly improved performance and task completion speed compared to the standard clinical information system. More than 1000 users and partners in the hospital tested and used VIEWER, reporting high satisfaction and expressed strong interest in incorporating VIEWER into their daily practice. DISCUSSION: VIEWER provides a cost-effective enhancement to the functionalities of standard clinical information systems, with evaluation offering valuable feedback for future improvements. CONCLUSION: VIEWER was developed to improve data accessibility and representation across various aspects of healthcare delivery, including population health management and patient monitoring. The deployment of VIEWER highlights the benefits of collaborative refinement in optimizing health informatics solutions for enhanced patient care. Tao Wang 0036, David Codling, Yamiko Joseph Msosa, Matthew Broadbent, Daisy Kornblum, Catherine Polling, Thomas Searle, Claire Delaney-Pope, Barbara Arroyo, Stuart MacLellan, Zoe Keddie, Mary Docherty, Angus Roberts, Robert Stewart 0002, Philip K. McGuire, Richard J. B. Dobson, Robert Harland |
J. Am. Medical Informatics Assoc. | 16 |
| 2026 | CSAI: Conditional Self-Attention Imputation for Healthcare Time-SeriesabstractWe introduce the Conditional Self-Attention Imputation (CSAI) model, a novel recurrent neural network architecture designed to address imputation challenges in multivariate time series derived from hospital electronic health records (EHRs). CSAI introduces key novelties specific to EHR data: a) attention-based hidden state initialisation to capture both long- and short-range temporal dependencies, b) domain-informed temporal decay to mimic clinical recording patterns, and c) a non-uniform masking strategy that models non-random missingness. Comprehensive evaluation across four EHR benchmark datasets demonstrates CSAI's effectiveness compared to state-of-the-art architectures in data restoration and downstream tasks. CSAI is integrated into PyPOTS, an open-source Python toolbox for partially observed time series. This work significantly advances the state of neural network imputation applied to EHRs by more closely aligning algorithmic imputation with clinical realities. Linglong Qian, Joseph Arul Raj, Hugh Logan Ellis, Yuezhou Zhang 0001, Tao Wang 0036, Richard J. B. Dobson, Zina M. Ibrahim |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Speech Reference Intervals: An Assessment of Feasibility in Depression Symptom Severity PredictionabstractMajor Depressive Disorder (MDD) is a prevalent mental disorder. Combining speech features and machine learning has promise for predicting MDD, but interpretability is crucial for clinical applications. Reference intervals (RIs) represent a typical range for a speech feature in a population. RIs could increase interpretability and help clinicians identify deviations from norms. They could also replace conventional speech features in machine learning models. However, no work has yet assessed the feasibility of speech RIs in MDD. We generated and compared RIs from three reference datasets varying in size, elicitation prompt, and health information. We then calculated deviations from each RI set for people with MDD to compare performance on a depression symptom severity prediction task. Our RI-based models trained with demographic data performed similarly to each other and equivalent models using conventional features or demographics only, demonstrating the value of RI-derived features. Lauren L. White, Ewan Carr, Judith Dineley, Catarina Botelho, Pauline Conde, Faith Matcham, Carolin Oetzmann, Amos Folarin, George Fairs, Agnes Norbury, Stefano Goria, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Alberto Abad, Isabel Trancoso, Nicholas Cummins |
INTERSPEECH | 14 |
| 2025 | How Deep is Your Guess? A Fresh Perspective on Deep Learning for Medical Time-Series ImputationabstractWe present a comprehensive analysis of deep learning approaches for Electronic Health Record (EHR) time-series imputation, examining how the interplay between architectural and framework design decisions gives rise to higher-level properties of a given deep imputer model and distinct biases towards complex data characteristics. Our investigation reveals the varying capabilities of deep imputers in capturing complex spatio-temporal dependencies within EHRs, and that the effectiveness of the model depends on how its combined biases align with the characteristics of the medical time series. Our experimental evaluation challenges common assumptions about model complexity, demonstrating that larger models do not necessarily improve performance. Rather, carefully designed architectures can better capture the complex patterns inherent in clinical data. The study highlights the need for imputation approaches that prioritise clinically meaningful data reconstruction over statistical accuracy. Our experiments further reveal up to 20% in variations of imputation performance based on preprocessing and implementation choices, emphasising the need for standardised benchmarking methodologies. Finally, we identify critical gaps between current deep imputation methods and medical requirements, highlighting the importance of integrating clinical insights to achieve more reliable imputation approaches for healthcare applications. Linglong Qian, Hugh Logan Ellis, Tao Wang 0036, Jun Wang 0121, Robin Mitra, Richard J. B. Dobson, Zina M. Ibrahim |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Longitudinal Modeling of Depression Shifts Using Speech and LanguageabstractSpeech analysis can provide a potential non-invasive and objective means of assessing and monitoring an individual’s mental health. Most studies to date have focused on cross-sectional analysis and have not explored the benefits of speech analysis as a longitudinal monitoring tool that can assist in the management of chronic conditions such as major depressive disorder (MDD). Objectively monitoring for shifts in depression symptom severity levels over time presents a notable challenge, which we address through an automated approach using longitudinal English and Spanish speech samples collected from a clinical population. We employ time–frequency representations and linguistic embeddings to enhance the early recognition of alterations in depression levels in individuals with MDD. We investigate the suitability of using siamese-based training for modeling these changes, intending to enable personalized and adaptive interventions. Paula Andrea Pérez-Toro, Judith Dineley, Agnieszka Kaczkowska, Pauline Conde, Yuezhou Zhang 0001, Faith Matcham, Sara Siddi, Josep Maria Haro, Stuart Bruce, Til Wykes, Raquel Bailón, Srinivasan Vairavan, Richard J. B. Dobson, Andreas K. Maier, Elmar Nöth, Juan Rafael Orozco-Arroyave, Vaibhav A. Narayan, Nicholas Cummins |
ICASSP | 13 |
| 2024 | Variability of speech timing features across repeated recordings: a comparison of open-source extraction techniquesabstractVariations in speech timing features have been reliably linked to symptoms of various health conditions, demonstrating clinical potential. However, replication challenges hinder their translation; extracted speech features are susceptible to methodological variations in the recording and processing pipeline. Investigating this, we compared exemplar timing features extracted via three different techniques from recordings of healthy speech. Our results show that features extracted via an intensity-based method differ from those produced by forced alignment. Different extraction methods also led to differing estimates of within-speaker feature variability over time in an analysis of recordings repeated systematically over three sessions in one day (n=26) and in one week (n=28). Our findings highlight the importance of feature extraction in study design and interpretation, and the need for consistent, accurate extraction techniques for clinical research. Index Terms: speech timing, feature extraction, reproducibility, longitudinal monitoring Judith Dineley, Ewan Carr, Lauren L. White, Catriona Lucas, Zahia Rahman, Tian Pan 0004, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 9 |
| 2024 | Real-Time Mobile Health Analytics and Interventions Pipeline to Detect Acute Events in COPD
Heet Sankesara, Yatharth Ranjan, Pauline Conde, Malik A. Althobiani, Zulqarnain Rashid, Akash Roy Choudhury, Callum L. Stewart, Yuezhou Zhang 0001, Joanna Porter, John R. Hurst, Richard J. B. Dobson, Amos Folarin |
MobiQuitous | 11 |
| 2024 | Predicting Future Disorders via Temporal Knowledge Graphs and Medical OntologiesabstractDespite the vast potential for insights and value present in Electronic Health Records (EHRs), it is challenging to fully leverage all the available information, particularly that contained in the free-text data written by clinicians describing the health status of patients. The utilization of Named Entity Recognition and Linking tools allows not only for the structuring of information contained within free-text data, but also for the integration with medical ontologies, which may prove highly beneficial for the analysis of patient medical histories with the aim of forecasting future medical outcomes, such as the diagnosis of a new disorder. In this paper, we propose MedTKG, a Temporal Knowledge Graph (TKG) framework that incorporates both the dynamic information of patient clinical histories and the static information of medical ontologies. The TKG is used to model a medical history as a series of snapshots at different points in time, effectively capturing the dynamic nature of the patient's health status, while a static graph is used to model the hierarchies of concepts extracted from domain ontologies. The proposed method aims to predict future disorders by identifying missing objects in the quadruple 〈s, r, ?, t 〉, where s and r denote the patient and the disorder relation type, respectively, and t is the timestamp of the query. The method is evaluated on clinical notes extracted from MIMIC-III and demonstrates the effectiveness of the TKG framework in predicting future disorders and of medical ontologies in improving its performance. Marco Postiglione, Daniel Bean, Zeljko Kraljevic, Richard J. B. Dobson, Vincenzo Moscato |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Classifying depression symptom severity: Assessment of speech representations in personalized and generalized machine learning modelsabstractThere is an urgent need for new methods that improve the management and treatment of Major Depressive Disorder (MDD). Speech has long been regarded as a promising digital marker in this regard, with many works highlighting that speech changes associated with MDD can be captured through machine learning models. Typically, findings are based on cross-sectional data, with little work exploring the advantages of personalization in building more robust and reliable models. This work assesses the strengths of different combinations of speech representations and machine learning models, in personalized and generalized settings in a two-class depression severity classification paradigm. Key results on a longitudinal dataset highlight the benefits of personalization. Our strongest performing model set-up utilized self-supervised learning features and convolutional neural network (CNN) and long short-term memory (LSTM) back-end. Edward L. Campbell, Judith Dineley, Pauline Conde, Faith Matcham, Katie M. White, Carolin Oetzmann, Sara Simblett, Stuart Bruce, Amos Folarin, Til Wykes, Srinivasan Vairavan, Richard J. B. Dobson, Laura Docío Fernández, Carmen García-Mateo, Vaibhav A. Narayan, Matthew Hotopf, Nicholas Cummins |
INTERSPEECH | 12 |
| 2023 | Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speechabstractSpeech is promising as an objective, convenient tool to monitor health remotely over time using mobile devices. Numerous paralinguistic features have been demonstrated to contain salient information related to an individual’s health. However, mobile device specification and acoustic environments vary widely, risking the reliability of the extracted features. In an initial step towards quantifying these effects, we report the variability of 13 exemplar paralinguistic features commonly reported in the speech-health literature and extracted from the speech of 42 healthy volunteers recorded consecutively in rooms with low and high reverberation with one budget and two higher-end smartphones, and a condenser microphone. Our results show reverberation has a clear effect on several features, in particular voice quality markers. They point to new research directions investigating how best to record and process in-the-wild speech for reliable longitudinal health state assessment. Judith Dineley, Ewan Carr, Faith Matcham, Johnny Downs, Richard J. B. Dobson, Thomas F. Quatieri, Nicholas Cummins |
INTERSPEECH | 5 |
| 2023 | DNAscan2: a versatile, scalable, and user-friendly analysis pipeline for human next-generation sequencing dataabstractSUMMARY: The current widespread adoption of next-generation sequencing (NGS) in all branches of basic research and clinical genetics fields means that users with highly variable informatics skills, computing facilities and application purposes need to process, analyse, and interpret NGS data. In this landscape, versatility, scalability, and user-friendliness are key characteristics for an NGS analysis software. We developed DNAscan2, a highly flexible, end-to-end pipeline for the analysis of NGS data, which (i) can be used for the detection of multiple variant types, including SNVs, small indels, transposable elements, short tandem repeats, and other large structural variants; (ii) covers all standard steps of NGS analysis, from quality control of raw data and genome alignment to variant calling, annotation, and generation of reports for the interpretation and prioritization of results; (iii) is highly adaptable as it can be deployed and run via either a graphic user interface for non-bioinformaticians and a command line tool for personal computer usage; (iv) is scalable as it can be executed in parallel as a Snakemake workflow, and; (v) is computationally efficient by minimizing RAM and CPU time requirements. AVAILABILITY AND IMPLEMENTATION: DNAscan2 is implemented in Python3 and is available at https://github.com/KHP-Informatics/DNAscanv2. Heather Marriott, Renata Kabiljo, Ahmad Al Khleifat, Richard J. B. Dobson, Ammar Al-Chalabi, Alfredo Iacoangeli |
Bioinform. | 4 |
| 2023 | Discharge summary hospital course summarisation of in patient Electronic Health Record text with clinical concept guided deep pre-trained Transformer models
Thomas Searle, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson |
J. Biomed. Informatics | 4 |
| 2023 | Trustworthy Data and AI Environments for Clinical Prediction: Application to Crisis-Risk in People With DepressionabstractDepression is a common mental health condition that often occurs in association with other chronic illnesses, and varies considerably in severity. Electronic Health Records (EHRs) contain rich information about a patient's medical history and can be used to train, test and maintain predictive models to support and improve patient care. This work evaluated the feasibility of implementing an environment for predicting mental health crisis among people living with depression based on both structured and unstructured EHRs. A large EHR from a mental health provider, Mersey Care, was pseudonymised and ingested into the Natural Language Processing (NLP) platform CogStack, allowing text content in binary clinical notes to be extracted. All unstructured clinical notes and summaries were semantically annotated by MedCAT and BioYODIE NLP services. Cases of crisis in patients with depression were then identified. Random forest models, gradient boosting trees, and Long Short-Term Memory (LSTM) networks, with varying feature arrangement, were trained to predict the occurrence of crisis. The results showed that all the prediction models can use a combination of structured and unstructured EHR information to predict crisis in patients with depression with good and useful accuracy. The LSTM network that was trained on a modified dataset with only 1000 most-important features from the random forest model with temporality showed the best performance with a mean AUC of 0.901 and a standard deviation of 0.006 using a training dataset and a mean AUC of 0.810 and 0.01 using a hold-out test dataset. Comparing the results from the technical evaluation with the views of psychiatrists shows that there are now opportunities to refine and integrate such prediction models into pragmatic point-of-care clinical decision support tools for supporting mental healthcare delivery. Yamiko Joseph Msosa, Arturas Grauslys, Tao Wang 0036, Iain E. Buchan, Paul Langan, Steven Foster, Michael Pearson, Amos Folarin, Angus Roberts, Simon Maskell, Richard J. B. Dobson, Cecil Kullu, Dennis Kehoe |
IEEE J. Biomed. Health Informatics | 13 |
| 2022 | Transforming and evaluating the UK Biobank to the OMOP Common Data Model for COVID-19 research and beyondabstractOBJECTIVE: The coronavirus disease 2019 (COVID-19) pandemic has demonstrated the value of real-world data for public health research. International federated analyses are crucial for informing policy makers. Common data models (CDMs) are critical for enabling these studies to be performed efficiently. Our objective was to convert the UK Biobank, a study of 500 000 participants with rich genetic and phenotypic data to the Observational Medical Outcomes Partnership (OMOP) CDM. MATERIALS AND METHODS: We converted UK Biobank data to OMOP CDM v. 5.3. We transformedparticipant research data on diseases collected at recruitment and electronic health records (EHRs) from primary care, hospitalizations, cancer registrations, and mortality from providers in England, Scotland, and Wales. We performed syntactic and semantic validations and compared comorbidities and risk factors between source and transformed data. RESULTS: We identified 502 505 participants (3086 with COVID-19) and transformed 690 fields (1 373 239 555 rows) to the OMOP CDM using 8 different controlled clinical terminologies and bespoke mappings. Specifically, we transformed self-reported noncancer illnesses 946 053 (83.91% of all source entries), cancers 37 802 (70.81%), medications 1 218 935 (88.25%), and prescriptions 864 788 (86.96%). In EHR, we transformed 13 028 182 (99.95%) hospital diagnoses, 6 465 399 (89.2%) procedures, 337 896 333 primary care diagnoses (CTV3, SNOMED-CT), 139 966 587 (98.74%) prescriptions (dm+d) and 77 127 (99.95%) deaths (ICD-10). We observed good concordance across demographic, risk factor, and comorbidity factors between source and transformed data. DISCUSSION AND CONCLUSION: Our study demonstrated that the OMOP CDM can be successfully leveraged to harmonize complex large-scale biobanked studies combining rich multimodal phenotypic data. Our study uncovered several challenges when transforming data from questionnaires to the OMOP CDM which require further research. The transformed UK Biobank resource is a valuable tool that can enable federated research, like COVID-19 studies. Václav Papez, Maxim Moinat, Erica A. Voss, Sofia Bazakou, Anne Van Winzum, Alessia Peviani, Stefan Payralbe, Elena Garcia Lara, Michael Kallfelz, Folkert W. Asselbergs, Daniel Prieto-Alhambra, Richard J. B. Dobson, Spiros C. Denaxas |
J. Am. Medical Informatics Assoc. | 12 |
| 2022 | Patient-centric characterization of multimorbidity trajectories in patients with severe mental illnesses: A temporal bipartite network modeling approachabstractMultimorbidity is a major factor contributing to increased mortality among people with severe mental illnesses (SMI). Previous studies either focus on estimating prevalence of a disease in a population without considering relationships between diseases or ignore heterogeneity of individual patients in examining disease progression by looking merely at aggregates across a whole cohort. Here, we present a temporal bipartite network model to jointly represent detailed information on both individual patients and diseases, which allows us to systematically characterize disease trajectories from both patient and disease centric perspectives. We apply this approach to a large set of longitudinal diagnostic records for patients with SMI collected through a data linkage between electronic health records from a large UK mental health hospital and English national hospital administrative database. We find that the resulting diagnosis networks show disassortative mixing by degree, suggesting that patients affected by a small number of diseases tend to suffer from prevalent diseases. Factors that determine the network structures include an individual's age, gender and ethnicity. Our analysis on network evolution further shows that patients and diseases become more interconnected over the illness duration of SMI, which is largely driven by the process that patients with similar attributes tend to suffer from the same conditions. Our analytic approach provides a guide for future patient-centric research on multimorbidity trajectories and contributes to achieving precision medicine. Tao Wang 0036, Rebecca Bendayan, Yamiko Msosa, Megan Pritchard, Angus Roberts, Robert Stewart 0002, Richard J. B. Dobson |
J. Biomed. Informatics | 7 |
| 2022 | Fitbeat: COVID-19 estimation based on wristband heart rate using a contrastive convolutional auto-encoder
Shuo Liu 0012, Jing Han 0010, Estela Laporta Puyal, Spyridon Kontaxis, Shaoxiong Sun, Patrick Locatelli, Judith Dineley, Florian B. Pokorny, Gloria Dalla Costa, Letizia Leocani, Ana Isabel Guerrero, Carlos Nos, Ana Zabalza, Per Soelberg Sørensen, Mathias Buron, Melinda Magyari, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Callum L. Stewart, Amos Folarin, Richard J. B. Dobson, Raquel Bailón, Srinivasan Vairavan, Nicholas Cummins, Vaibhav A. Narayan, Matthew Hotopf, Giancarlo Comi, Björn W. Schuller |
Pattern Recognit. | 22 |
| 2022 | A Knowledge Distillation Ensemble Framework for Predicting Short- and Long-Term Hospitalization Outcomes From Electronic Health Records DataabstractThe ability to perform accurate prognosis is crucial for proactive clinical decision making, informed resource management and personalised care. Existing outcome prediction models suffer from a low recall of infrequent positive outcomes. We present a highly-scalable and robust machine learning framework to automatically predict adversity represented by mortality and ICU admission and readmission from time-series of vital signs and laboratory results obtained within the first 24 hours of hospital admission. The stacked ensemble platform comprises two components: a) an unsupervised LSTM Autoencoder that learns an optimal representation of the time-series, using it to differentiate the less frequent patterns which conclude with an adverse event from the majority patterns that do not, and b) a gradient boosting model, which relies on the constructed representation to refine prediction by incorporating static features. The model is used to assess a patient's risk of adversity and provides visual justifications of its prediction. Results of three case studies show that the model outperforms existing platforms in ICU and general ward settings, achieving average Precision-Recall Areas Under the Curve (PR-AUCs) of 0.891 (95% CI: 0.878-0.939) for mortality and 0.908 (95% CI: 0.870-0.935) in predicting ICU admission and readmission. Zina M. Ibrahim, Daniel Bean, Thomas Searle, Linglong Qian, Honghan Wu, Anthony Shek, Zeljko Kraljevic, James Galloway, Sam Norton, James T. Teo, Richard J. B. Dobson |
IEEE J. Biomed. Health Informatics | 11 |
| 2021 | Remote Smartphone-Based Speech Collection: Acceptance and Barriers in Individuals with Major Depressive DisorderabstractThe ease of in-the-wild speech recording using smartphones has sparked considerable interest in the combined application of speech, remote measurement technology (RMT) and advanced analytics as a research and healthcare tool. For this to be realised, the acceptability of remote speech collection to the user must be established, in addition to feasibility from an analytical perspective. To understand the acceptance, facilitators, and barriers of smartphone-based speech recording, we invited 384 individuals with major depressive disorder (MDD) from the Remote Assessment of Disease and Relapse - Central Nervous System (RADAR-CNS) research programme in Spain and the UK to complete a survey on their experiences recording their speech. In this analysis, we demonstrate that study participants were more comfortable completing a scripted speech task than a free speech task. For both speech tasks, we found depression severity and country to be significant predictors of comfort. Not seeing smartphone notifications of the scheduled speech tasks, low mood and forgetfulness were the most commonly reported obstacles to providing speech recordings. Judith Dineley, Grace Lavelle, Daniel Leightley, Faith Matcham, Sara Siddi, Maria Teresa Peñarrubia-María, Katie M. White, Alina Ivan, Carolin Oetzmann, Sara Simblett, Erin Dawe-Lane, Stuart Bruce, Daniel Stahl, Yatharth Ranjan, Zulqarnain Rashid, Pauline Conde, Amos Folarin, Josep Maria Haro, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Björn W. Schuller, Nicholas Cummins |
Interspeech | 20 |
| 2021 | Multi-domain clinical natural language processing with MedCAT: The Medical Concept Annotation Toolkit
Zeljko Kraljevic, Thomas Searle, Anthony Shek, Lukasz Roguski, Kawsar Noor, Daniel Bean, Aurelie Mascio, Leilei Zhu, Amos Folarin, Angus Roberts, Rebecca Bendayan, Mark P. Richardson, Robert Stewart 0002, Anoop D. Shah, Wai Keong Wong, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson |
Artif. Intell. Medicine | 18 |
| 2021 | Ensemble learning for poor prognosis predictions: A case study on SARS-CoV-2abstractOBJECTIVE: Risk prediction models are widely used to inform evidence-based clinical decision making. However, few models developed from single cohorts can perform consistently well at population level where diverse prognoses exist (such as the SARS-CoV-2 [severe acute respiratory syndrome coronavirus 2] pandemic). This study aims at tackling this challenge by synergizing prediction models from the literature using ensemble learning. MATERIALS AND METHODS: In this study, we selected and reimplemented 7 prediction models for COVID-19 (coronavirus disease 2019) that were derived from diverse cohorts and used different implementation techniques. A novel ensemble learning framework was proposed to synergize them for realizing personalized predictions for individual patients. Four diverse international cohorts (2 from the United Kingdom and 2 from China; N = 5394) were used to validate all 8 models on discrimination, calibration, and clinical usefulness. RESULTS: Results showed that individual prediction models could perform well on some cohorts while poorly on others. Conversely, the ensemble model achieved the best performances consistently on all metrics quantifying discrimination, calibration, and clinical usefulness. Performance disparities were observed in cohorts from the 2 countries: all models achieved better performances on the China cohorts. DISCUSSION: When individual models were learned from complementary cohorts, the synergized model had the potential to achieve better performances than any individual model. Results indicate that blood parameters and physiological measurements might have better predictive powers when collected early, which remains to be confirmed by further studies. CONCLUSIONS: Combining a diverse set of individual prediction models, the ensemble method can synergize a robust and well-performing model by choosing the most competent ones for individual patients. Honghan Wu, Andreas Karwath, Zina M. Ibrahim, Kevin Dhaliwal, Daniel Bean, Victor Roth Cardoso, Kezhi Li, James T. Teo, Amitava Banerjee, Fang Gao-Smith, Tony Whitehouse, Tonny Veenith, Georgios V. Gkoutos, Richard J. B. Dobson, Bruce Guthrie |
J. Am. Medical Informatics Assoc. | 20 |
| 2021 | Estimating redundancy in clinical text
Thomas Searle, Zina M. Ibrahim, James T. Teo, Richard J. B. Dobson |
J. Biomed. Informatics | 4 |
| 2020 | Modeling Rare Interactions in Time Series Data Through Qualitative Change: Application to Outcome Prediction in Intensive Care UnitsabstractMany areas of research are characterised by the deluge of large-scale highly-dimensional time-series data. However, using the data available for prediction and decision making is hampered by the current lag in our ability to uncover and quantify true interactions that explain the outcomes. We are interested in areas such as intensive care medicine, which are characterised by i) continuous monitoring of multivariate variables and non-uniform sampling of data streams, ii) the outcomes are generally governed by interactions between a small set of rare events, iii) these interactions are not necessarily definable by specific values (or value ranges) of a given group of variables, but rather, by the deviations of these values from the normal state recorded over time, iv) the need to explain the predictions made by the model. Here, while numerous data mining models have been formulated for outcome prediction, they are unable to explain their predictions. We present a model for uncovering interactions with the highest likelihood of generating the outcomes seen from highly-dimensional time series data. Interactions among variables are represented by a relational graph structure, which relies on qualitative abstractions to overcome non-uniform sampling and to capture the semantics of the interactions corresponding to the changes and deviations from normality of variables of interest over time. Using the assumption that similar templates of small interactions are responsible for the outcomes (as prevalent in the medical domains), we reformulate the discovery task to retrieve the most-likely templates from the data. Experiments on sepsis prediction using real Intensive Care Unit (ICU) data demonstrates that the discovered interaction templates are semantically meaningful within the domain, and using them as features in a prediction task produces a superior performance than when using the raw values of the predictors. Zina M. Ibrahim, Honghan Wu, Richard J. B. Dobson |
ECAI | 3 |
| 2020 | Comparing Natural Language Processing Techniques for Alzheimer's Dementia Prediction in Spontaneous SpeechabstractAlzheimer's Dementia (AD) is an incurable, debilitating, and progressive neurodegenerative condition that affects cognitive function. Early diagnosis is important as therapeutics can delay progression and give those diagnosed vital time. Developing models that analyse spontaneous speech could eventually provide an efficient diagnostic modality for earlier diagnosis of AD. The Alzheimer's Dementia Recognition through Spontaneous Speech task offers acoustically pre-processed and balanced datasets for the classification and prediction of AD and associated phenotypes through the modelling of spontaneous speech. We exclusively analyse the supplied textual transcripts of the spontaneous speech dataset, building and comparing performance across numerous models for the classification of AD vs controls and the prediction of Mental Mini State Exam scores. We rigorously train and evaluate Support Vector Machines (SVMs), Gradient Boosting Decision Trees (GBDT), and Conditional Random Fields (CRFs) alongside deep learning Transformer based models. We find our top performing models to be a simple Term Frequency-Inverse Document Frequency (TF-IDF) vectoriser as input into a SVM model and a pre-trained Transformer based model `DistilBERT' when used as an embedding layer into simple linear models. We demonstrate test set scores of 0.81-0.82 across classification metrics and a RMSE of 4.58. Thomas Searle, Zina M. Ibrahim, Richard J. B. Dobson |
INTERSPEECH | 3 |
| 2020 | On classifying sepsis heterogeneity in the ICU: insight using machine learningabstractOBJECTIVES: Current machine learning models aiming to predict sepsis from electronic health records (EHR) do not account 20 for the heterogeneity of the condition despite its emerging importance in prognosis and treatment. This work demonstrates the added value of stratifying the types of organ dysfunction observed in patients who develop sepsis in the intensive care unit (ICU) in improving the ability to recognize patients at risk of sepsis from their EHR data. MATERIALS AND METHODS: Using an ICU dataset of 13 728 records, we identify clinically significant sepsis subpopulations with distinct organ dysfunction patterns. We perform classification experiments with random forest, gradient boost trees, and support vector machines, using the identified subpopulations to distinguish patients who develop sepsis in the ICU from those who do not. RESULTS: The classification results show that features selected using sepsis subpopulations as background knowledge yield a superior performance in distinguishing septic from non-septic patients regardless of the classification model used. The improved performance is especially pronounced in specificity, which is a current bottleneck in sepsis prediction machine learning models. CONCLUSION: Our findings can steer machine learning efforts toward more personalized models for complex conditions including sepsis. Zina M. Ibrahim, Honghan Wu, Ahmed Hamoud, Lukas Stappen, Richard J. B. Dobson, Andrea Agarossi |
J. Am. Medical Informatics Assoc. | 5 |
| 2020 | Natural Language Processing for Mimicking Clinical Trial Recruitment in Critical Care: A Semi-Automated Simulation Based on the LeoPARDS TrialabstractClinical trials often fail to recruit an adequate number of appropriate patients. Identifying eligible trial participants is resource-intensive when relying on manual review of clinical notes, particularly in critical care settings where the time window is short. Automated review of electronic health records (EHR) may help, but much of the information is in free text rather than a computable form. We applied natural language processing (NLP) to free text EHR data using the CogStack platform to simulate recruitment into the LeoPARDS study, a clinical trial aiming to reduce organ dysfunction in septic shock. We applied an algorithm to identify eligible patients using a moving 1-hour time window, and compared patients identified by our approach with those actually screened and recruited for the trial, for the time period that data were available. We manually reviewed records of a random sample of patients identified by the algorithm but not screened in the original trial. Our method identified 376 patients, including 34 patients with EHR data available who were actually recruited to LeoPARDS in our centre. The sensitivity of CogStack for identifying patients screened was 90% (95% CI 85%, 93%). Of the 203 patients identified by both manual screening and CogStack, the index date matched in 95 (47%) and CogStack was earlier in 94 (47%). In conclusion, analysis of EHR data using NLP could effectively replicate recruitment in a critical care trial, and identify some eligible patients at an earlier stage, potentially improving trial recruitment if implemented in real time. Hegler Tissot, Anoop D. Shah, David Brealey, Steve K. Harris, Ruth Agbakoba, Amos Folarin, Luis Romao, Lukasz Roguski, Richard J. B. Dobson, Folkert W. Asselbergs |
IEEE J. Biomed. Health Informatics | 9 |
| 2019 | DNAscan: personal computer compatible NGS analysis, annotation and visualisationabstractBACKGROUND: Next Generation Sequencing (NGS) is a commonly used technology for studying the genetic basis of biological processes and it underpins the aspirations of precision medicine. However, there are significant challenges when dealing with NGS data. Firstly, a huge number of bioinformatics tools for a wide range of uses exist, therefore it is challenging to design an analysis pipeline. Secondly, NGS analysis is computationally intensive, requiring expensive infrastructure, and many medical and research centres do not have adequate high performance computing facilities and cloud computing is not always an option due to privacy and ownership issues. Finally, the interpretation of the results is not trivial and most available pipelines lack the utilities to favour this crucial step. RESULTS: We have therefore developed a fast and efficient bioinformatics pipeline that allows for the analysis of DNA sequencing data, while requiring little computational effort and memory usage. DNAscan can analyse a whole exome sequencing sample in 1 h and a 40x whole genome sequencing sample in 13 h, on a midrange computer. The pipeline can look for single nucleotide variants, small indels, structural variants, repeat expansions and viral genetic material (or any other organism). Its results are annotated using a customisable variety of databases and are available for an on-the-fly visualisation with a local deployment of the gene.iobio platform. DNAscan is implemented in Python. Its code and documentation are available on GitHub: https://github.com/KHP-Informatics/DNAscan . Instructions for an easy and fast deployment with Docker and Singularity are also provided on GitHub. CONCLUSIONS: DNAscan is an extremely fast and computationally efficient pipeline for analysis, visualization and interpretation of NGS data. It is designed to provide a powerful and easy-to-use tool for applications in biomedical research and diagnostic medicine, at minimal computational cost. Its comprehensive approach will maximise the potential audience of users, bringing such analyses within the reach of non-specialist laboratories, and those from centres with limited funding available. Alfredo Iacoangeli, Ahmad Al Khleifat, William Sproviero, Aleksey Shatunov, A. R. Jones, S. L. Morgan, Alan Pittman, Richard J. B. Dobson, S. J. Newhouse, Ammar Al-Chalabi |
BMC Bioinform. | 8 |
| 2019 | UK phenomics platform for developing and validating electronic health record phenotypes: CALIBERabstractOBJECTIVE: Electronic health records (EHRs) are a rich source of information on human diseases, but the information is variably structured, fragmented, curated using different coding systems, and collected for purposes other than medical research. We describe an approach for developing, validating, and sharing reproducible phenotypes from national structured EHR in the United Kingdom with applications for translational research. MATERIALS AND METHODS: We implemented a rule-based phenotyping framework, with up to 6 approaches of validation. We applied our framework to a sample of 15 million individuals in a national EHR data source (population-based primary care, all ages) linked to hospitalization and death records in England. Data comprised continuous measurements (for example, blood pressure; medication information; coded diagnoses, symptoms, procedures, and referrals), recorded using 5 controlled clinical terminologies: (1) read (primary care, subset of SNOMED-CT [Systematized Nomenclature of Medicine Clinical Terms]), (2) International Classification of Diseases-Ninth Revision and Tenth Revision (secondary care diagnoses and cause of mortality), (3) Office of Population Censuses and Surveys Classification of Surgical Operations and Procedures, Fourth Revision (hospital surgical procedures), and (4) DM+D prescription codes. RESULTS: Using the CALIBER phenotyping framework, we created algorithms for 51 diseases, syndromes, biomarkers, and lifestyle risk factors and provide up to 6 validation approaches. The EHR phenotypes are curated in the open-access CALIBER Portal (https://www.caliberresearch.org/portal) and have been used by 40 national and international research groups in 60 peer-reviewed publications. CONCLUSIONS: We describe a UK EHR phenomics approach within the CALIBER EHR data platform with initial evidence of validity and use, as an important step toward international use of UK EHR data for health research. Spiros C. Denaxas, Arturo Gonzalez-Izquierdo, Kenan Direk, Natalie K. Fitzpatrick, Ghazaleh Fatemifar, Amitava Banerjee, Richard J. B. Dobson, Laurence J. Howe, Valerie Kuan, R. Tom Lumbers, Laura Pasea, Riyaz S. Patel, Anoop D. Shah, Aroon D. Hingorani, Cathie Sudlow, Harry Hemingway |
J. Am. Medical Informatics Assoc. | 7 |
| 2018 | SemEHR: A general-purpose semantic search system to surface semantic data from clinical notes for tailored care, trial recruitment, and clinical researchabstractObjective: Unlocking the data contained within both structured and unstructured components of electronic health records (EHRs) has the potential to provide a step change in data available for secondary research use, generation of actionable medical insights, hospital management, and trial recruitment. To achieve this, we implemented SemEHR, an open source semantic search and analytics tool for EHRs. Methods: SemEHR implements a generic information extraction (IE) and retrieval infrastructure by identifying contextualized mentions of a wide range of biomedical concepts within EHRs. Natural language processing annotations are further assembled at the patient level and extended with EHR-specific knowledge to generate a timeline for each patient. The semantic data are serviced via ontology-based search and analytics interfaces. Results: SemEHR has been deployed at a number of UK hospitals, including the Clinical Record Interactive Search, an anonymized replica of the EHR of the UK South London and Maudsley National Health Service Foundation Trust, one of Europe's largest providers of mental health services. In 2 Clinical Record Interactive Search-based studies, SemEHR achieved 93% (hepatitis C) and 99% (HIV) F-measure results in identifying true positive patients. At King's College Hospital in London, as part of the CogStack program (github.com/cogstack), SemEHR is being used to recruit patients into the UK Department of Health 100 000 Genomes Project (genomicsengland.co.uk). The validation study suggests that the tool can validate previously recruited cases and is very fast at searching phenotypes; time for recruitment criteria checking was reduced from days to minutes. Validated on open intensive care EHR data, Medical Information Mart for Intensive Care III, the vital signs extracted by SemEHR can achieve around 97% accuracy. Conclusion: Results from the multiple case studies demonstrate SemEHR's efficiency: weeks or months of work can be done within hours or minutes in some cases. SemEHR provides a more comprehensive view of patients, bringing in more and unexpected insight compared to study-oriented bespoke IE systems. SemEHR is open source, available at https://github.com/CogStack/SemEHR. Honghan Wu, Giulia Toti, Katherine Morley, Zina M. Ibrahim, Amos Folarin, Richard G. Jackson, Ismail Emre Kartoglu, Asha Agrawal, Clive Stringer, Darren Gale, Genevieve Gorrell, Angus Roberts, Matthew T. M. Broadbent, Robert Stewart 0002, Richard J. B. Dobson |
J. Am. Medical Informatics Assoc. | 15 |
| 2015 | On Energy-Efficient Computations With Advice
Hans-Joachim Böckenhauer, Richard J. B. Dobson, Sacha Krug, Kathleen Steinhöfel |
COCOON | 2 |
| 2014 | TextHunter - A User Friendly Tool for Extracting Generic Concepts from Free Text in Clinical Research
Richard G. Jackson, Michael Ball 0004, Rashmi Patel, Richard D. Hayes, Richard J. B. Dobson, Robert Stewart 0002 |
AMIA | 5 |
| 2013 | Detecting epistasis in the presence of linkage disequilibrium: A focused comparisonabstractWe present results from a comparison of three epistasis-detection tools using large-scale simulated genetic data: SNPHarvester, SNPRuler and Ambience. The tools were chosen based on their merits to be representative of the state of the art of epistasis detection. We design and conduct experiments to test the performance of the methods in detecting interacting loci or their proxies in linkage disequilibrium (LD) tagged regions, in datasets containing simulated 2,3 and 4-way epistatic interactions. The results show that SNPHarvester is the fastest while Ambience is the most robust. Moreover, SNPRuler provides the best power, specially with higher-level interactions, but cannot scale-up to larger datasets. Zina M. Ibrahim, Stephen Newhouse, Richard J. B. Dobson |
CIBCB | 3 |
| 2006 | Predicting deleterious nsSNPs: an analysis of sequence and structural attributesabstractBACKGROUND: There has been an explosion in the number of single nucleotide polymorphisms (SNPs) within public databases. In this study we focused on non-synonymous protein coding single nucleotide polymorphisms (nsSNPs), some associated with disease and others which are thought to be neutral. We describe the distribution of both types of nsSNPs using structural and sequence based features and assess the relative value of these attributes as predictors of function using machine learning methods. We also address the common problem of balance within machine learning methods and show the effect of imbalance on nsSNP function prediction. We show that nsSNP function prediction can be significantly improved by 100% undersampling of the majority class. The learnt rules were then applied to make predictions of function on all nsSNPs within Ensembl. RESULTS: The measure of prediction success is greatly affected by the level of imbalance in the training dataset. We found the balanced dataset that included all attributes produced the best prediction. The performance as measured by the Matthews correlation coefficient (MCC) varied between 0.49 and 0.25 depending on the imbalance. As previously observed, the degree of sequence conservation at the nsSNP position is the single most useful attribute. In addition to conservation, structural predictions made using a balanced dataset can be of value. CONCLUSION: The predictions for all nsSNPs within Ensembl, based on a balanced dataset using all attributes, are available as a DAS annotation. Instructions for adding the track to Ensembl are at http://www.brightstudy.ac.uk/das_help.html. Richard J. B. Dobson, Patricia B. Munroe, Mark J. Caulfield, Mansoor A. S. Saqi |
BMC Bioinform. | 1 |