VLDB 2026 Research / reviewers in the wild / expert
Tingting Zhu 0001
dblp:29/7666-1
· DBLP profile ↗
38ranked-venue papers
3as first author
27since 2021 · last 2026
0000-0002-1552-5630ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A perspective on individualized treatment effects estimation from time-series health dataabstractOBJECTIVES: The objective of this study is to provide an overview of the current landscape of individualized treatment effects (ITE) estimation, specifically focusing on methodologies proposed for time-series electronic health records (EHRs). We aim to identify gaps in the literature, discuss challenges, and propose future research directions to advance the field of personalized medicine. MATERIALS AND METHODS: We conducted a comprehensive literature review to identify and analyze relevant works on ITE estimation for time-series data. The review focused on theoretical assumptions, types of treatment settings, and computational frameworks employed in the existing literature. RESULTS: The literature reveals a growing body of work on ITE estimation for tabular data, while methodologies specific to time-series EHRs are limited. We summarize and discuss the latest advancements, including the types of models proposed, the theoretical foundations, and the computational approaches used. DISCUSSION: The limitations and challenges of current ITE estimation methods for time-series data are discussed, including the lack of standardized evaluation metrics and the need for more diverse and representative datasets. We also highlight considerations and potential biases that may arise in personalized treatment effect estimation. CONCLUSION: This work provides a comprehensive overview of ITE estimation for time-series EHR data, offering insights into the current state of the field and identifying future research directions. By addressing the limitations and challenges, we hope to encourage further exploration and innovation in this exciting and under-studied area of personalized medicine. Ghadeer O. Ghosheh, Moritz Gögl, Tingting Zhu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2026 | DySurv: dynamic deep learning model for survival analysis with conditional variational inferenceabstractOBJECTIVE: Machine learning applications for longitudinal electronic health records often forecast the risk of events at fixed time points, whereas survival analysis achieves dynamic risk prediction by estimating time-to-event distributions. Here, we propose a novel conditional variational autoencoder-based method, DySurv, which uses a combination of static and longitudinal measurements from electronic health records to estimate the individual risk of death dynamically. MATERIALS AND METHODS: DySurv directly estimates the cumulative risk incidence function without making any parametric assumptions on the underlying stochastic process of the time-to-event. We evaluate DySurv on 6 time-to-event benchmark datasets in healthcare, as well as 2 real-world intensive care unit (ICU) electronic health records (EHR) datasets extracted from the eICU Collaborative Research (eICU) and the Medical Information Mart for Intensive Care database (MIMIC-IV). RESULTS: DySurv outperforms other existing statistical and deep learning approaches to time-to-event analysis across concordance and other metrics. It achieves time-dependent concordance of over 60% in the eICU case. It is also over 12% more accurate and 22% more sensitive than in-use ICU scores like Acute Physiology and Chronic Health Evaluation (APACHE) and Sequential Organ Failure Assessment (SOFA) scores. The predictive capacity of DySurv is consistent and the survival estimates remain disentangled across different datasets. DISCUSSION: Our interdisciplinary framework successfully incorporates deep learning, survival analysis, and intensive care to create a novel method for time-to-event prediction from longitudinal health records. We test our method on several held-out test sets from a variety of healthcare datasets and compare it to existing in-use clinical risk scoring benchmarks. CONCLUSION: While our method leverages non-parametric extensions to deep learning-guided estimations of the survival distribution, further deep learning paradigms could be explored. Munib Mesinovic, Peter J. Watkinson, Tingting Zhu 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2026 | Real-World Classification of Student Stress and Fatigue Using Wearable PPG RecordingsabstractWearable-based affective computing offers a promising solution for monitoring and managing stress and fatigue in adolescent student populations. This can help to support student well-being without increasing reliance on mobile phones by providing on-device insights into stress and fatigue levels. To do so, the processing and classification methods must be lightweight enough to be executed in real-time on wearables. This study outlines the findings of the implementation of a student-informed wearable and mobile app, Wellby. Students tested this for one month while completing routine photoplethysmography (PPG) recordings and check-ins on their perceived levels of stress and fatigue. This study proposes a lightweight processing pipeline, intended for wearable-based deployment, while examining its classification performance on real-world student PPG data from Wellby. The pipeline performs denoising, fixed noise elimination, and peak detection to calculate time-domain heart rate variability (HRV) metrics. It was first evaluated on public datasets, the Wearable Stress and Affect Detection (WESAD) dataset and the AKTIVES dataset, achieving an area under the receiver operating characteristic curve (AUC-ROC) of up to 91.58% for stress classification on WESAD and 76.61% on AKTIVES. In the Wellby dataset, the adapted processing pipeline achieved an AUC-ROC of 77.02% for stress classification and 71.58% for fatigue classification using only time-domain HRV features. Furthermore, the inclusion of a signal quality metric and baseline well-being questionnaires improved the AUC-ROC for stress classification to 91.60% in the best performing model. These findings demonstrate the potential for wearables to implement real-time affective computing, providing timely feedback to students in real-world settings based on PPG and contextual data. The code used in this study is available on GitHub [https://github.com/j-laiti/PPG-affect-classification]. Justin Laiti, Yu Liu 0016, Pádraic J. Dunne, Elaine Byrne, Tingting Zhu 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | AnchorInv: Few-Shot Class-Incremental Learning of Physiological Signals via Feature Space-Guided InversionabstractDeep learning models have demonstrated exceptional performance in a variety of real-world applications. These successes are often attributed to strong base models that can generalize to novel tasks with limited supporting data while keeping prior knowledge intact. However, these impressive results are based on the availability of a large amount of high-quality data, which is often lacking in specialized biomedical applications. In such fields, models are often developed with limited data that arrive incrementally with novel categories. This requires the model to adapt to new information while preserving existing knowledge. Few-Shot Class-Incremental Learning (FSCIL) methods offer a promising approach to addressing these challenges, but they also depend on strong base models that face the same aforementioned limitations. To overcome these constraints, we propose AnchorInv following the straightforward and efficient buffer-replay strategy. Instead of selecting and storing raw data, AnchorInv generates synthetic samples guided by anchor points in the feature space. This approach protects privacy and regularizes the model for adaptation. When evaluated on three public physiological time series datasets, AnchorInv exhibits efficient knowledge forgetting prevention and improved adaptation to novel classes, surpassing state-of-the-art baselines. Chenqi Li, Boyan Gao, Gabriel Davis Jones, Timothy Denison, Tingting Zhu 0001 |
AAAI | 5 |
| 2025 | ProtoEHR: Hierarchical Prototype Learning for EHR-based Healthcare PredictionsabstractDigital healthcare systems have enabled the collection of mass healthcare data in electronic healthcare records (EHRs), allowing artificial intelligence solutions for various healthcare prediction tasks. However, existing studies often focus on isolated components of EHR data, limiting their predictive performance and interpretability. To address this gap, we propose ProtoEHR, an interpretable hierarchical prototype learning framework that fully exploits the rich, multi-level structure of EHR data to enhance healthcare predictions. More specifically, ProtoEHR models relationships within and across three hierarchical levels of EHRs: medical codes, hospital visits, and patients. We first leverage large language models to extract semantic relationships among medical codes and construct a medical knowledge graph as the knowledge source. Building on this, we design a hierarchical representation learning framework that captures contextualized representations across three levels, while incorporating prototype information within each level to capture intrinsic similarities and improve generalization. To perform a comprehensive assessment, we evaluate ProtoEHR in two public datasets on five clinically significant tasks, including prediction of mortality, prediction of readmission, prediction of length of stay, drug recommendation, and prediction of phenotype. The results demonstrate the ability of ProtoEHR to make accurate, robust, and interpretable predictions compared to baselines in the literature. Furthermore, ProtoEHR offers interpretable insights on code, visit, and patient levels to aid in healthcare prediction. Zi Cai, Yu Liu 0016, Zhiyao Luo, Tingting Zhu 0001 |
CIKM | 4 |
| 2025 | SurvUnc: A Meta-Model Based Uncertainty Quantification Framework for Survival AnalysisabstractSurvival analysis, which estimates the probability of event occurrence over time from censored data, is fundamental in numerous real-world applications, particularly in high-stakes domains such as healthcare and risk assessment. Despite advances in numerous survival models, quantifying the uncertainty of predictions from these models remains underexplored and challenging. The lack of reliable uncertainty quantification limits the interpretability and trustworthiness of survival models, hindering their adoption in clinical decision-making and other sensitive applications. To bridge this gap, in this work, we introduce SurvUnc, a novel meta-model based framework for post-hoc uncertainty quantification for survival models. SurvUnc introduces an anchor-based learning strategy that integrates concordance knowledge into meta-model optimization, leveraging pairwise ranking performance to estimate uncertainty effectively. Notably, our framework is model-agnostic, ensuring compatibility with any survival model without requiring modifications to its architecture or access to its internal parameters. Especially, we design a comprehensive evaluation pipeline tailored to this critical yet overlooked problem. Through extensive experiments on four publicly available benchmarking datasets and five representative survival models, we demonstrate the superiority of SurvUnc across multiple evaluation scenarios, including selective prediction, misprediction detection, and out-of-domain detection. Our results highlight the effectiveness of SurvUnc in enhancing model interpretability and reliability, paving the way for more trustworthy survival predictions in real-world Yu Liu 0016, Weiyao Tao, Tong Xia, Simon Knight 0005, Tingting Zhu 0001 |
KDD (2) | 5 |
| 2025 | DoseSurv: Predicting Personalized Survival Outcomes under Continuous-Valued TreatmentsabstractEstimating heterogeneous treatment effects (HTEs) of continuous-valued interventions on survival, that is, time-to-event (TTE) outcomes, is crucial in various fields, notably in clinical decision-making and in driving the advancement of next-generation clinical trials. However, while HTE estimation for continuous-valued (i.e., dosage-dependent) interventions and for TTE outcomes have been separately explored, their combined application remains largely overlooked in the machine learning literature. We propose DoseSurv, a varying-coefficient network designed to estimate HTEs for different dosage-dependent and non-dosage treatment options from TTE data. DoseSurv uses radial basis functions to model continuity in dose-response relationships and learns balanced representations to address covariate shifts arising in HTE estimation from observational TTE data. We present experiments across various treatment scenarios on both simulated and real-world data, demonstrating DoseSurv's superior performance over existing baseline models. Moritz Gögl, Yu Liu 0016, Christopher Yau, Peter J. Watkinson, Tingting Zhu 0001 |
NeurIPS | 5 |
| 2025 | Benchmarking Large Language Models in Evidence-Based MedicineabstractEvidence-based medicine (EBM) represents a paradigm of providing patient care grounded in the most current and rigorously evaluated research. Recent advances in large language models (LLMs) offer a potential solution to transform EBM by automating labor-intensive tasks and thereby improving the efficiency of clinical decision-making. This study explores integrating LLMs into the key stages in EBM, evaluating their ability across evidence retrieval (PICO extraction, biomedical question answering), synthesis (summarizing randomized controlled trials), and dissemination (medical text simplification). We conducted a comparative analysis of seven LLMs, including both proprietary and open-source models, as well as those fine-tuned on medical corpora. Specifically, we benchmarked the performance of various LLMs on each EBM task under zero-shot settings as baselines, and employed prompting techniques, including in-context learning, chain-of-thought reasoning, and knowledge-guided prompting to enhance their capabilities. Our extensive experiments revealed the strengths of LLMs, such as remarkable understanding capabilities even in zero-shot settings, strong summarization skills, and effective knowledge transfer via prompting. Promoting strategies such as knowledge-guided prompting proved highly effective (e.g., improving the performance of GPT-4 by 13.10% over zero-shot in PICO extraction). However, the experiments also showed limitations, with LLM performance falling well below state-of-the-art baselines like PubMedBERT in handling named entity recognition tasks. Moreover, human evaluation revealed persisting challenges with factual inconsistencies and domain inaccuracies, underscoring the need for rigorous quality control before clinical application. This study provides insights into enhancing EBM using LLMs while highlighting critical areas for further research. Jin Li 0067, Yiyan Deng, Yu Tian 0002, Jingsong Li 0001, Tingting Zhu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Uncertainty-Inspired Multi-Task Learning in Arbitrary Scenarios of ECG MonitoringabstractAs the scenarios for electrocardiogram (ECG) monitoring become increasingly diverse, particularly with the development of wearable ECG, the influence of ambiguous factors in diagnosis has been amplified. Reliable ECG information must be extracted from abundant noises and confusing artifacts. To address this issue, we suggest an uncertainty-inspired model for beat-level diagnosis (UI-Beat). The base architecture of UI-Beat separates heartbeat localization and event diagnosis in two branches to address the problem of heterogeneous data sources. To disentangle the epistemic and aleatoric uncertainty within one stage in a deterministic neural network, we propose a new method derived from uncertainty formulation and realize it by introducing the class-biased transformation. Then the disentangled uncertainty can be utilized to screen out noise and identify ambiguous heartbeat synchronously. The results indicate that UI-Beat can significantly improve the performance of noise detection (from 91.60% to 97.50% for real-world noise detection and from 61.40% to 82.41% for real-world artifact detection). For multi-lead ECG analysis, UI-Beat is approaching the performance upper bound in heartbeat localization (only 15 false positives and 9 false negatives out of the 175,907 heartbeats in the INCART database) and achieving a significant performance improvement in heartbeat classification through uncertainty-based cross-lead fusion compared to single-lead prediction and other state-of-the-art methods (an average improvement of 14.28% for detecting heartbeats of S and 3.37% for detecting heartbeats of V). Considering the characteristic of one-stage ECG analysis within one model, it is suggested that the proposed UI-Beat has the potential to be employed as a general model for arbitrary scenarios of ECG monitoring, with the capacity to remove unusableepisodes, and realize heartbeat-level diagnosis with confidence provided. Xingyao Wang 0001, Hongxiang Gao, Caiyun Ma, Tingting Zhu 0001, Feng Yang 0011, Chengyu Liu 0001, Huazhu Fu |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Position: Reinforcement Learning in Dynamic Treatment Regimes Needs Critical ReexaminationabstractIn the rapidly changing healthcare landscape, the implementation of offline reinforcement learning (RL) in dynamic treatment regimes (DTRs) presents a mix of unprecedented opportunities and challenges. This position paper offers a critical examination of the current status of offline RL in the context of DTRs. We argue for a reassessment of applying RL in DTRs, citing concerns such as inconsistent and potentially inconclusive evaluation metrics, the absence of naive and supervised learning baselines, and the diverse choice of RL formulation in existing research. Through a case study with more than 17,000 evaluation experiments using a publicly available Sepsis dataset, we demonstrate that the performance of RL algorithms can significantly vary with changes in evaluation metrics and Markov Decision Process (MDP) formulations. Surprisingly, it is observed that in some instances, RL algorithms can be surpassed by random baselines subjected to policy evaluation methods and reward design. This calls for more careful policy evaluation and algorithm development in future DTR works. Additionally, we discussed potential enhancements toward more reliable development of RL-based dynamic treatment regimes and invited further discussion within the community. Code is available at https://github.com/GilesLuo/ReassessDTR. Zhiyao Luo, Yangchen Pan, Peter J. Watkinson, Tingting Zhu 0001 |
ICML | 4 |
| 2024 | Temporal dynamics unleashed: Elevating variational graph attentionabstractThis research introduces the Variational Graph Attention Dynamics (VarGATDyn), addressing the complexities of dynamic graph representation learning, where existing models, tailored for static graphs, prove inadequate. VarGATDyn melds attention mechanisms with a Markovian assumption to surpass the challenges of maintaining temporal consistency and the extensive dataset requirements typical of RNN-based frameworks. It harnesses the strengths of the Variational Graph Auto-Encoder (VGAE) framework, Graph Attention Networks (GAT), and Gaussian Mixture Models (GMM) to adeptly navigate the temporal and structural intricacies of dynamic graphs. Through the strategic application of GMMs, the model handles multimodal patterns, thereby rectifying misalignments between prior and estimated posterior distributions. An innovative multiple-learning methodology bolsters the model's adaptability, leading to an encompassing and effective learning process. Empirical tests underscore VarGATDyn's dominance in dynamic link prediction across various datasets, highlighting its proficiency in capturing multimodal distributions and temporal dynamics. Soheila Molaei, Ghazaleh Niknam, Ghadeer O. Ghosheh, Vinod Kumar Chauhan, Hadi Zare 0001, Tingting Zhu 0001, Shirui Pan, David A. Clifton |
Knowl. Based Syst. | 6 |
| 2024 | AutoNet-Generated Deep Layer-Wise Convex Networks for ECG ClassificationabstractThe design of neural networks typically involves trial-and-error, a time-consuming process for obtaining an optimal architecture, even for experienced researchers. Additionally, it is widely accepted that loss functions of deep neural networks are generally non-convex with respect to the parameters to be optimised. We propose the Layer-wise Convex Theorem to ensure that the loss is convex with respect to the parameters of a given layer, achieved by constraining each layer to be an overdetermined system of non-linear equations. Based on this theorem, we developed an end-to-end algorithm (the AutoNet) to automatically generate layer-wise convex networks (LCNs) for any given training set. We then demonstrate the performance of the AutoNet-generated LCNs (AutoNet-LCNs) compared to state-of-the-art models on three electrocardiogram (ECG) classification benchmark datasets, with further validation on two non-ECG benchmark datasets for more general tasks. The AutoNet-LCN was able to find networks customised for each dataset without manual fine-tuning under 2 GPU-hours, and the resulting networks outperformed the state-of-the-art models with fewer than 5% parameters on all the above five benchmark datasets. The efficiency and robustness of the AutoNet-LCN markedly reduce model discovery costs and enable efficient training of deep learning models in resource-constrained settings. Yanting Shen, Tingting Zhu 0001, Xinshao Wang, Lei A. Clifton, Zhengming Chen 0003, Robert Clarke, David A. Clifton |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Student Loss: Towards the Probability Assumption in Inaccurate SupervisionabstractNoisy labels are often encountered in datasets, but learning with them is challenging. Although natural discrepancies between clean and mislabeled samples in a noisy category exist, most techniques in this field still gather them indiscriminately, which leads to their performances being partially robust. In this paper, we reveal both empirically and theoretically that the learning robustness can be improved by assuming deep features with the same labels follow a student distribution, resulting in a more intuitive method called student loss. By embedding the student distribution and exploiting the sharpness of its curve, our method is naturally data-selective and can offer extra strength to resist mislabeled samples. This ability makes clean samples aggregate tightly in the center, while mislabeled samples scatter, even if they share the same label. Additionally, we employ the metric learning strategy and develop a large-margin student (LT) loss for better capability. It should be noted that our approach is the first work that adopts the prior probability assumption in feature representation to decrease the contributions of mislabeled samples. This strategy can enhance various losses to join the student loss family, even if they have been robust losses. Experiments demonstrate that our approach is more effective in inaccurate supervision. Enhanced LT losses significantly outperform various state-of-the-art methods in most cases. Even huge improvements of over 50% can be obtained under some conditions. Shuo Zhang 0030, Jianqing Li 0002, Hamido Fujita, Yuwen Li 0002, Dengbao Wang, Tingting Zhu 0001, Min-Ling Zhang, Chengyu Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Intelligent Electrocardiogram Acquisition Via Ubiquitous Photoplethysmography MonitoringabstractRecent advances in machine learning, particularly deep neural network architectures, have shown substantial promise in classifying and predicting cardiac abnormalities from electrocardiogram (ECG) data. Such data are rich in information content, typically in morphology and timing, due to the close correlation between cardiac function and the ECG. However, the ECG is usually not measured ubiquitously in a passive manner from consumer devices, and generally requires 'active' sampling whereby the user prompts a device to take an ECG measurement. Conversely, photoplethysmography (PPG) data are typically measured passively by consumer devices, and therefore available for long-period monitoring and suitable in duration for identifying transient cardiac events. However, classifying or predicting cardiac abnormalities from the PPG is very difficult, because it is a peripherally-measured signal. Hence, the use of the PPG for predictive inference is often limited to deriving physiological parameters (heart rate, breathing rate, etc.) or for obvious abnormalities in cardiac timing, such as atrial fibrillation/flutter ("palpitations"). This work aims to combine the best of both worlds: using continuously-monitored, near-ubiquitous PPG to identify periods of sufficient abnormality in the PPG such that prompting the user to take an ECG would be informative of cardiac risk. We propose a dual-convolutional-attention network (DCA-Net) to achieve this ECG-based PPG classification. With DCA-Net, we prove the plausibility of this concept on MIMIC Waveform Database with high performance level (AUROC 0.9 and AUPRC 0.7) and receive satisfactory result when testing the model on an independent dataset (AUROC 0.7 and AUPRC 0.6) which it is not perfectly-matched to the MIMIC dataset. Zhangdaihong Liu, Tingting Zhu 0001, Yuan-Ting Zhang, David A. Clifton |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Incremental Trainable Parameter Selection in Deep Neural NetworksabstractThis article explores the utilization of the effective degree-of-freedom (DoF) of a deep learning model to regularize its stochastic gradient descent (SGD)-based training. The effective DoF of a deep learning model is defined only by a subset of its total parameters. This subset is highly responsive or sensitive toward the training loss, and its cardinality can be used to govern the effective DoF of a model during training. To this aim, the incremental trainable parameter selection (ITPS) algorithm is introduced in this article. The proposed ITPS algorithm acts as a wrapper over SGD and incrementally selects the parameters for updation that exhibit the maximum sensitivity toward the training loss. Hence, it gradually increases the DoF of the model during training. In ideal cases, the proposed algorithm arrives at a model configuration (i.e., DoF) optimum for the task at hand. This whole process results in a regularization-like behavior induced by a gradual increment of the DoF. Since the selection and updation of parameters is a function of the training loss, the proposed algorithm can be seen as a task and data-dependent regularization mechanism. This article exhibits the general utility of ITPS by evaluating it on various prominent neural network architectures such as CNNs, transformers, recurrent neural networks (RNNs), and multilayer perceptrons. These models are trained for image classification and healthcare tasks using the publicly available CIFAR-10, SLT-10, and MIMIC-III datasets. Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Tingting Zhu 0001, David A. Clifton |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Adversarial De-confounding in Individualised Treatment Effects EstimationabstractObservational studies have recently received significant attention from the machine learning community due to the increasingly available non-experimental observational data and the limitations of the experimental studies, such as considerable cost, impracticality, small and less representative sample sizes, etc. In observational studies, de-confounding is a fundamental problem of individualised treatment effects (ITE) estimation. This paper proposes disentangled representations with adversarial training to selectively balance the confounders in the binary treatment setting for the ITE estimation. The adversarial training of treatment policy selectively encourages treatment-agnostic balanced representations for the confounders and helps to estimate the ITE in the observational studies via counterfactual inference. Empirical results on synthetic and real-world datasets, with varying degrees of confounding, prove that our proposed approach improves the state-of-the-art methods in achieving lower error in the ITE estimation. Vinod Kumar Chauhan, Soheila Molaei, Marzia Hoque Tania, Anshul Thakur, Tingting Zhu 0001, David A. Clifton |
AISTATS | 5 |
| 2023 | Learning personalized preference: A segmentation strategy under consumer sparse data
Tingting Zhu 0001, Ye-Zheng Liu 0001 |
Expert Syst. Appl. | 1 |
| 2023 | DyVGRNN: DYnamic mixture Variational Graph Recurrent Neural Networks
Ghazaleh Niknam, Soheila Molaei, Hadi Zare 0001, Shirui Pan, Mahdi Jalili, Tingting Zhu 0001, David A. Clifton |
Neural Networks | 6 |
| 2022 | Learning of Cluster-based Feature Importance for Electronic Health Record Time-seriesabstractThe recent availability of Electronic Health Records (EHR) has allowed for the development of algorithms predicting inpatient risk of deterioration and trajectory evolution. However, prediction of disease progression with EHR is challenging since these data are sparse, heterogeneous, multi-dimensional, and multi-modal time-series. As such, clustering is regularly used to identify similar groups within the patient cohort to improve prediction. Current models have shown some success in obtaining cluster representations of patient trajectories. However, they i) fail to obtain clinical interpretability for each cluster, and ii) struggle to learn meaningful cluster numbers in the context of imbalanced distribution of disease outcomes. We propose a supervised deep learning model to cluster EHR data based on the identification of clinically understandable phenotypes with regard to both outcome prediction and patient trajectory. We introduce novel loss functions to address the problems of class imbalance and cluster collapse, and furthermore propose a feature-time attention mechanism to identify cluster-based phenotype importance across time and feature dimensions. We tested our model in two datasets corresponding to distinct medical settings. Our model yielded added interpretability to cluster formation and outperformed benchmarks by at least 4% in relevant metrics. Henrique Aguiar, Mauro D. Santos, Peter J. Watkinson, Tingting Zhu 0001 |
ICML | 4 |
| 2022 | SoQal: Selective Oracle Questioning for Consistency Based Active Learning of Cardiac SignalsabstractClinical settings are often characterized by abundant unlabelled data and limited labelled data. This is typically driven by the high burden placed on oracles (e.g., physicians) to provide annotations. One way to mitigate this burden is via active learning (AL) which involves the (a) acquisition and (b) annotation of informative unlabelled instances. Whereas previous work addresses either one of these elements independently, we propose an AL framework that addresses both. For acquisition, we propose Bayesian Active Learning by Consistency (BALC), a sub-framework which perturbs both instances and network parameters and quantifies changes in the network output probability distribution. For annotation, we propose SoQal, a sub-framework that dynamically determines whether, for each acquired unlabelled instance, to request a label from an oracle or to pseudo-label it instead. We show that BALC can outperform start-of-the-art acquisition functions such as BALD, and SoQal outperforms baseline methods even in the presence of a noisy oracle. Dani Kiyasseh, Tingting Zhu 0001, David A. Clifton |
ICML | 2 |
| 2022 | Topic discovery from short reviews based on data enhancementabstractWith the rapid development of social media and mobile Internet, short reviews, such as Weibo and Twitter, have exploded online. Discovering topics from short reviews is significant for many practical applications. It can effectively not only identify users’ attitudes and emotions but also enhance customer satisfaction and shopping experience. Because reviews are relatively short, the sparsity of reviews considerably restricts the quality of topic discovery. To improve the efficiency of topic discovery, we introduce the concept of data enhancement and strengthen the data in sentences and words in short reviews based on the weight of importance. We then propose a topic model for reviews to topic discovery based on data enhancement (shorted as DE-LDA). We verify the rationality and feasibility of DE-LDA on real datasets. Results show that the proposed method outperforms benchmarks in topic discovery and also has better clustering effects. Tingting Zhu 0001, Ye-Zheng Liu 0001, Jianshan Sun, Chunhua Sun |
Intell. Data Anal. | 1 |
| 2021 | Prediction of monetary penalties for data protection cases in multiple languagesabstractAs the use of personal data becomes further entrenched in the function of societal interaction, the regulation of such data continues to grow as an important area of law. Nevertheless, it is unfortunately the case that data protection authorities have limited resources to address an increasing number of investigations. The leveraging of appropriate data-driven models, coupled with the automation of decision making, has the potential to help in such circumstances. In this paper, we evaluate machine learning models in the literature (such as Support Vector Machine (SVM), Random Forest, and Multinomial Naive Bayes (MNB) classifiers) for natural language processing in order to predict whether a monetary penalty was levied based on a description of case facts. We tested these models on a novel data set collected from the data protection authority of Macao across the three languages (i.e., Chinese, English, and Portuguese). Our experimental results show that the machine learning models provide the necessary predictability in order to automate the evaluation of data protection cases. In particular, SVM has consistent performance across three languages and achieving an AUROC of 0.725, 0.762, and 0.748 for Chinese, English, and Portuguese, respectively. We further evaluated the interpretability of the results independently for each of the languages and found that the salient texts that were identified are shared across the three languages. Aaron Ceross, Tingting Zhu 0001 |
ICAIL | 2 |
| 2021 | CLOCS: Contrastive Learning of Cardiac Signals Across Space, Time, and PatientsabstractThe healthcare industry generates troves of unlabelled physiological data. This data can be exploited via contrastive learning, a self-supervised pre-training method that encourages representations of instances to be similar to one another. We propose a family of contrastive learning methods, CLOCS, that encourages representations across space, time, \textit{and} patients to be similar to one another. We show that CLOCS consistently outperforms the state-of-the-art methods, BYOL and SimCLR, when performing a linear evaluation of, and fine-tuning on, downstream tasks. We also show that CLOCS achieves strong generalization performance with only 25% of labelled training data. Furthermore, our training procedure naturally generates patient-specific representations that can be used to quantify patient-similarity. Dani Kiyasseh, Tingting Zhu 0001, David A. Clifton |
ICML | 2 |
| 2021 | CROCS: Clustering and Retrieval of Cardiac Signals Based on Patient Disease Class, Sex, and AgeabstractThe process of manually searching for relevant instances in, and extracting information from, clinical databases underpin a multitude of clinical tasks. Such tasks include disease diagnosis, clinical trial recruitment, and continuing medical education. This manual search-and-extract process, however, has been hampered by the growth of large-scale clinical databases and the increased prevalence of unlabelled instances. To address this challenge, we propose a supervised contrastive learning framework, CROCS, where representations of cardiac signals associated with a set of patient-specific attributes (e.g., disease class, sex, age) are attracted to learnable embeddings entitled clinical prototypes. We exploit such prototypes for both the clustering and retrieval of unlabelled cardiac signals based on multiple patient attributes. We show that CROCS outperforms the state-of-the-art method, DTC, when clustering and also retrieves relevant cardiac signals from a large database. We also show that clinical prototypes adopt a semantically meaningful arrangement based on patient attributes and thus confer a high degree of interpretability. Dani Kiyasseh, Tingting Zhu 0001, David A. Clifton |
NeurIPS | 2 |
| 2021 | DeepMI: Deep multi-lead ECG fusion for identifying myocardial infarction and its occurrence-time
Girmaw Abebe, Hamza A. Javed, Komminist Weldemariam, Jiyan Chen, Tingting Zhu 0001 |
Artif. Intell. Medicine | 7 |
| 2021 | Stroke risk prediction using machine learning: a prospective cohort study of 0.5 million Chinese adultsabstractOBJECTIVE: To compare Cox models, machine learning (ML), and ensemble models combining both approaches, for prediction of stroke risk in a prospective study of Chinese adults. MATERIALS AND METHODS: We evaluated models for stroke risk at varying intervals of follow-up (<9 years, 0-3 years, 3-6 years, 6-9 years) in 503 842 adults without prior history of stroke recruited from 10 areas in China in 2004-2008. Inputs included sociodemographic factors, diet, medical history, physical activity, and physical measurements. We compared discrimination and calibration of Cox regression, logistic regression, support vector machines, random survival forests, gradient boosted trees (GBT), and multilayer perceptrons, benchmarking performance against the 2017 Framingham Stroke Risk Profile. We then developed an ensemble approach to identify individuals at high risk of stroke (>10% predicted 9-yr stroke risk) by selectively applying either a GBT or Cox model based on individual-level characteristics. RESULTS: For 9-yr stroke risk prediction, GBT provided the best discrimination (AUROC: 0.833 in men, 0.836 in women) and calibration, with consistent results in each interval of follow-up. The ensemble approach yielded incrementally higher accuracy (men: 76%, women: 80%), specificity (men: 76%, women: 81%), and positive predictive value (men: 26%, women: 24%) compared to any of the single-model approaches. DISCUSSION AND CONCLUSION: Among several approaches, an ensemble model combining both GBT and Cox models achieved the best performance for identifying individuals at high risk of stroke in a contemporary study of Chinese adults. The results highlight the potential value of expanding the use of ML in clinical practice. Matthew Chun, Robert Clarke, Benjamin J. Cairns, David A. Clifton, Derrick Bennett, Pei Pei, Canqing Yu, Zhengming Chen 0003, Tingting Zhu 0001 |
J. Am. Medical Informatics Assoc. | 14 |
| 2021 | Hospital Admission Location Prediction via Deep Interpretable Networks for the Year-Round Improvement of Emergency Patient CareabstractOBJECTIVE: This paper presents a deep learning method of predicting where in a hospital emergency patients will be admitted after being triaged in the Emergency Department (ED). Such a prediction will allow for the preparation of bed space in the hospital for timely care and admission of the patient as well as allocation of resource to the relevant departments, including during periods of increased demand arising from seasonal peaks in infections. METHODS: The problem is posed as a multi-class classification into seven separate ward types. A novel deep learning training strategy was created that combines learning via curriculum and a multi-armed bandit to exploit this curriculum post-initial training. RESULTS: We successfully predict the initial hospital admission location with area-under-receiver-operating-curve (AUROC) ranging between 0.60 to 0.78 for the individual wards and an overall maximum accuracy of 52% where chance corresponds to 14% for this seven-class setting. Our proposed network was able to interpret which features drove the predictions using a 'network saliency' term added to the network loss function. CONCLUSION: We have proven that prediction of location of admission in hospital for emergency patients is possible using information from triage in ED. We have also shown that there are certain tell-tale tests which indicate what space of the hospital a patient will use. SIGNIFICANCE: It is hoped that this predictor will be of value to healthcare institutions by allowing for the planning of resource and bed space ahead of the need for it. This in turn should speed up the provision of care for the patient and allow flow of patients out of the ED thereby improving patient flow and the quality of care for the remaining patients within the ED. Rasheed El-Bouri, David Eyre 0001, Peter J. Watkinson, Tingting Zhu 0001, David A. Clifton |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Student-Teacher Curriculum Learning via Reinforcement Learning: Predicting Hospital Inpatient Admission LocationabstractAccurate and reliable prediction of hospital admission location is important due to resource-constraints and space availability in a clinical setting, particularly when dealing with patients who come from the emergency department. In this work we propose a student-teacher network via reinforcement learning to deal with this specific problem. A representation of the weights of the student network is treated as the state and is fed as an input to the teacher network. The teacher network’s action is to select the most appropriate batch of data to train the student network on from a training set sorted according to entropy. By validating on three datasets, not only do we show that our approach outperforms state-of-the-art methods on tabular data and performs competitively on image recognition, but also that novel curricula are learned by the teacher network. We demonstrate experimentally that the teacher network can actively learn about the student network and guide it to achieve better performance than if trained alone. Rasheed El-Bouri, David Eyre 0001, Peter J. Watkinson, Tingting Zhu 0001, David A. Clifton |
ICML | 4 |
| 2020 | Discriminant Knowledge Extraction from Electrocardiograms for Automated Diagnosis of Myocardial Infarction
Girmaw Abebe, Komminist Weldemariam, Hamza A. Javed, Jiyan Chen, Tingting Zhu 0001 |
PKAW | 7 |
| 2020 | PlethAugment: GAN-Based PPG Augmentation for Medical Diagnosis in Low-Resource SettingsabstractThe paucity of physiological time-series data collected from low-resource clinical settings limits the capabilities of modern machine learning algorithms in achieving high performance. Such performance is further hindered by class imbalance; datasets where a diagnosis is much more common than others. To overcome these two issues at low-cost while preserving privacy, data augmentation methods can be employed. In the time domain, the traditional method of time-warping could alter the underlying data distribution with detrimental consequences. This is prominent when dealing with physiological conditions that influence the frequency components of data. In this paper, we propose PlethAugment; three different conditional generative adversarial networks (CGANs) with an adapted diversity term for the generation of pathological photoplethysmogram (PPG) signals in order to boost medical classification performance. To evaluate and compare the GANs, we introduce a novel metric-agnostic method; the synthetic generalization curve. We validate this approach on two proprietary and two public datasets representing a diverse set of medical conditions. Compared to training on non-augmented class-balanced datasets, training on augmented datasets leads to an improvement of the AUROC by up to 29% when using cross validation. This illustrates the potential of the proposed CGANs to significantly improve classification performance. Dani Kiyasseh, Girmaw Abebe, Nhan Le Nguyen Thanh, Le Van Tan, Louise Thwaites, Tingting Zhu 0001, David A. Clifton |
IEEE J. Biomed. Health Informatics | 6 |
| 2020 | Deep Interpretable Early Warning System for the Detection of Clinical DeteriorationabstractAssessment of physiological instability preceding adverse events on hospital wards has been previously investigated through clinical early warning score systems. Early warning scores are simple to use yet they consider data as independent and identically distributed random variables. Deep learning applications are able to learn from sequential data, however they lack interpretability and are thus difficult to deploy in clinical settings. We propose the 'Deep Early Warning System' (DEWS), an interpretable end-to-end deep learning model that interpolates temporal data and predicts the probability of an adverse event, defined as the composite outcome of cardiac arrest, mortality or unplanned ICU admission. The model was developed and validated using routinely collected vital signs of patients admitted to the the Oxford University Hospitals between 21st March 2014 and 31st March 2018. We extracted 45 314 vital-sign measurements as a balanced training set and 359 481 vital-sign measurements as an imbalanced testing set to mimic a real-life setting of emergency admissions. DEWS achieved superior accuracy than the state-of-the-art that is currently implemented in clinical settings, the National Early Warning Score, in terms of the overall area under the receiver operating characteristic curve (AUROC) (0.880 vs. 0.866) and when evaluated independently for each of the three outcomes. Our attention-based architecture was able to recognize 'historical' trends in the data that are most correlated with the predicted probability. With high sensitivity, improved clinical utility and increased interpretability, our model can be easily deployed in clinical settings to supplement existing EWS systems. Farah Shamout, Tingting Zhu 0001, Pulkit Sharma, Peter J. Watkinson, David A. Clifton |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Multi-Modal Diagnosis of Infectious Diseases in the Developing WorldabstractIn low and middle income countries, infectious diseases continue to have a significant impact, particularly amongst the poorest in society. Tetanus and hand foot and mouth disease (HFMD) are two such diseases and, in both, death is associated with autonomic nervous system dysfunction (ANSD). Currently, photoplethysmogram or electrocardiogram monitoring is used to detect deterioration in these patients, however expensive clinical monitors are often required. In this study, we employ low-cost and mobile wearable devices to collect patient vital signs unobtrusively; and we develop machine learning algorithms for automatic and rapid triage of patients that provide efficient use of clinical resources. Existing methods are mainly dependent on the prior detection of clinical features with limited exploitation of multi-modal physiological data. Moreover, the latest developments in deep learning (e.g. cross-domain transfer learning) have not been sufficiently applied for infectious disease diagnosis. In this paper, we present a fusion of multi-modal physiological data to predict the severity of ANSD with a hierarchy of resource-aware decision making. First, an on-site triage process is performed using a simple classifier. Second, personalised longitudinal modelling is employed that takes the previous states of the patient into consideration. We have also employed a spectrogram representation of the physiological waveforms to exploit existing networks for cross-domain transfer learning, which avoids the laborious and data intensive process of training a network from scratch. Results show that the proposed framework has promising potential in supporting severity grading of infectious diseases in low-resources settings, such as in the developing world. Girmaw Abebe, Hamza A. Javed, Nhan Le Nguyen Thanh, Ha Thi Hai Duong, Le Van Tan, Louise Thwaites, David A. Clifton, Tingting Zhu 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2019 | DeepAMR for predicting co-occurrent resistance of Mycobacterium tuberculosisabstractMOTIVATION: Resistance co-occurrence within first-line anti-tuberculosis (TB) drugs is a common phenomenon. Existing methods based on genetic data analysis of Mycobacterium tuberculosis (MTB) have been able to predict resistance of MTB to individual drugs, but have not considered the resistance co-occurrence and cannot capture latent structure of genomic data that corresponds to lineages. RESULTS: We used a large cohort of TB patients from 16 countries across six continents where whole-genome sequences for each isolate and associated phenotype to anti-TB drugs were obtained using drug susceptibility testing recommended by the World Health Organization. We then proposed an end-to-end multi-task model with deep denoising auto-encoder (DeepAMR) for multiple drug classification and developed DeepAMR_cluster, a clustering variant based on DeepAMR, for learning clusters in latent space of the data. The results showed that DeepAMR outperformed baseline model and four machine learning models with mean AUROC from 94.4% to 98.7% for predicting resistance to four first-line drugs [i.e. isoniazid (INH), ethambutol (EMB), rifampicin (RIF), pyrazinamide (PZA)], multi-drug resistant TB (MDR-TB) and pan-susceptible TB (PANS-TB: MTB that is susceptible to all four first-line anti-TB drugs). In the case of INH, EMB, PZA and MDR-TB, DeepAMR achieved its best mean sensitivity of 94.3%, 91.5%, 87.3% and 96.3%, respectively. While in the case of RIF and PANS-TB, it generated 94.2% and 92.2% sensitivity, which were lower than baseline model by 0.7% and 1.9%, respectively. t-SNE visualization shows that DeepAMR_cluster captures lineage-related clusters in the latent space. AVAILABILITY AND IMPLEMENTATION: The details of source code are provided at http://www.robots.ox.ac.uk/∼davidc/code.php. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yang Yang 0125, Timothy M. Walker, A. Sarah Walker, Daniel J. Wilson 0001, Tim E. A. Peto, Derrick W. Crook, Farah Shamout, Tingting Zhu 0001, David A. Clifton |
Bioinform. | 9 |
| 2019 | Service matchmaking for Internet of Things based on probabilistic topic model
Ye-Zheng Liu 0001, Tingting Zhu 0001, Yuan-Chun Jiang, Xiao Liu 0004 |
Future Gener. Comput. Syst. | 2 |
| 2019 | Identifying social roles using heterogeneous features in online social networksabstractRole analysis plays an important role when exploring social media and knowledge‐sharing platforms for designing marking strategies. However, current methods in role analysis have overlooked content generated by users (e.g., posts) in social media and hence focus more on user behavior analysis. The user‐generated content is very important for characterizing users. In this paper, we propose a novel method which integrates both user behavior and posted content by users to identify roles in online social networks. The proposed method models a role as a joint distribution of Gaussian distribution and multinomial distribution, which represent user behavioral feature and content feature respectively. The proposed method can be used to determine the number of roles concerned automatically. The experimental results show that the proposed method can be used to identify various roles more effectively and to get more insights on such characteristics. Ye-Zheng Liu 0001, Jianshan Sun, Thushari P. Silva, Yuan-Chun Jiang, Tingting Zhu 0001 |
J. Assoc. Inf. Sci. Technol. | 6 |
| 2019 | Unsupervised Bayesian Inference to Fuse Biosignal Sensory Estimates for Personalizing CareabstractThe role of sensing technologies, such as wearables, in delivering precision care is becoming widely acceptable. Given the very large quantities of sensor data that rapidly accumulate, there is a need to employ automated algorithms to label biosignal sensor data. In many real-life clinical applications, no such expert labels are available, and algorithms for processing sensor data must be relied upon, without access to the "ground truth." It is therefore extremely difficult to choose which algorithms to trust or discard at any point in time, where different algorithms may be optimal for different patients, or even for different points in time for the same patient. We propose two fully Bayesian approaches for fusing labels from independent and potentially correlated annotators (i.e., algorithms or, where available, experts). These are generative models to aggregate labels (i.e., the outputs of the algorithms, such as identified ECG morphology) in an unsupervised manner, to estimate jointly the assumed bias and precision of each algorithm without access to the ground truth. The latter fused estimate may then be used to infer the underlying ground truth. For the first time in the biomedical context, we show that modeling correlations between annotators, and fusing information concerning task difficulty (such as the estimated quality of the sensor data), improve these estimates with respect to commonly employed strategies in the literature. Also, we adopt a strongly Bayesian approach to inference using Gibbs sampling to improve estimates over the existing state of the art. We present results from applying the proposed pair of models to simulated and two publicly available biomedical datasets, to demonstrate proof-of-principle. We show that our proposed models outperform all existing approaches recreated from the literature. We also show that the proposed methods are robust when dealing with missing values (as often occurs in real-life biomedical applications), and that they are suitably efficient for use in real-time applications, thereby providing the basis for the reliable use of sensors for personalizing the care of the individual. Tingting Zhu 0001, Marco A. F. Pimentel, Gari D. Clifford, David A. Clifton |
IEEE J. Biomed. Health Informatics | 1 |
| 2018 | Machine learning for classifying tuberculosis drug-resistance from DNA sequencing dataabstractMotivation: Correct and rapid determination of Mycobacterium tuberculosis (MTB) resistance against available tuberculosis (TB) drugs is essential for the control and management of TB. Conventional molecular diagnostic test assumes that the presence of any well-studied single nucleotide polymorphisms is sufficient to cause resistance, which yields low sensitivity for resistance classification. Summary: Given the availability of DNA sequencing data from MTB, we developed machine learning models for a cohort of 1839 UK bacterial isolates to classify MTB resistance against eight anti-TB drugs (isoniazid, rifampicin, ethambutol, pyrazinamide, ciprofloxacin, moxifloxacin, ofloxacin, streptomycin) and to classify multi-drug resistance. Results: Compared to previous rules-based approach, the sensitivities from the best-performing models increased by 2-4% for isoniazid, rifampicin and ethambutol to 97% (P < 0.01), respectively; for ciprofloxacin and multi-drug resistant TB, they increased to 96%. For moxifloxacin and ofloxacin, sensitivities increased by 12 and 15% from 83 and 81% based on existing known resistance alleles to 95% and 96% (P < 0.01), respectively. Particularly, our models improved sensitivities compared to the previous rules-based approach by 15 and 24% to 84 and 87% for pyrazinamide and streptomycin (P < 0.01), respectively. The best-performing models increase the area-under-the-ROC curve by 10% for pyrazinamide and streptomycin (P < 0.01), and 4-8% for other drugs (P < 0.01). Availability and implementation: The details of source code are provided at http://www.robots.ox.ac.uk/~davidc/code.php. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Yang Yang 0125, Katherine E. Niehaus, Timothy M. Walker, Zamin Iqbal, A. Sarah Walker, Daniel J. Wilson 0001, Tim E. A. Peto, Derrick W. Crook, E. Grace Smith, Tingting Zhu 0001, David A. Clifton |
Bioinform. | 10 |
| 2018 | A crowdsourcing-based topic model for service matchmaking in Internet of Things
Ye-Zheng Liu 0001, Jianshan Sun, Yuan-Chun Jiang, Jianmin He, Tingting Zhu 0001, Chunhua Sun |
Future Gener. Comput. Syst. | 6 |