Puja Myles

dblp:232/3009 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-8976-890XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Predicting Performance Drift in AI Models of Healthcare Without Ground Truth Labels
Ylenia Rotalinti, Puja Myles, Allan Tucker
IDA (1)2
2023 The Impact of Bias on Drift Detection in AI Health Software
Asal Khoshravan Azar, Barbara Draghi, Ylenia Rotalinti, Puja Myles, Allan Tucker
AIME4
2023 Creating Synthetic Geospatial Patient Data to Mimic Real Data Whilst Preserving Privacy: *2022 35th International Symposium on Computer-Based Medical Systems (CBMS)
abstract
Synthetic Individual-Level Geospatial Data (SIL-GSD) offers a number of advantages in Spatial Epidemiology when compared to census data or surveys conducted on regional or global levels. The use of SILGSD could bring a new dimension to the study of the patterns and causes of diseases in a particular location while minimizing the risk of patient identity disclosure, especially for rare conditions. Additionally, it could help in building and monitoring regional machine learning models, improving the quality and effectiveness of local healthcare services. Finally, SILGSD could help in controlling the spread and causes of diseases by studying disease movement across areas through the travelling patterns of populations. To our knowledge, no synthetic health records data containing synthesised geographic locations for patients has been published for research purposes so far. Therefore, in this paper we explore generating SILGSD by allocating synthetic patients to general practices (healthcare providers) in the UK using the demographics and prevalence of health conditions in each practice. The assigned general practice locations can be used as proxies for patient locations due to people being registered to their nearest practice from home. We use high-fidelity synthetic primary care patients from the Clinical Practice Research Datalink (CPRD) and allocate them to England's general practices (GPs), using the publicly available GP health conditions statistics from the Quality and Outcomes Framework (QOF). The allocation relies on similarities between patients in different locations without using real location information for the patients. We demonstrate that the Allocation Data is able to accurately mimic the real health conditions distribution in the general practices and also preserves the underlying distribution of the original primary care patients data from CPRD (Gold Standard).
Dima R. Alattal, Zhenchen Wang, Puja Myles, Allan Tucker
CBMS3
2023 Privacy Assessment of Synthetic Patient Data
abstract
In this paper, we quantify the privacy gain of synthetic patient data drawn from two generative models, MST and PrivBayes, which is based on real anonymized primary care patient data. This evaluation is implemented for two types of inference attacks, namely membership and attribute inference attacks using a new toolbox, TAPAS. The aim is to quantitatively evaluate the privacy gain of each attack where these two differentially private generators and different threat models are used with a focus on black-box knowledge. The evaluation that was carried out in this paper demonstrates that vulnerabilities of synthetic patient data depend on the different attack scenarios, threat models, and algorithms used to generate the synthetic patient data. It was shown empirically that although the synthetic patient data achieved high privacy gain in most attack scenarios, it does not behave uniformly against adversarial attacks, and some records and outliers remain vulnerable depending on the attack scenario. Moreover, it was shown that the PrivBayes generator is the more robust generator in comparison to MST in terms of the privacy-preservation of synthetic data.
Ferdoos Hossein Nezhad, Ylenia Rotalinti, Puja Myles, Allan Tucker
CBMS3
2021 Evaluating a Longitudinal Synthetic Data Generator using Real World Data
abstract
Synthetic data offer a number of advantages over using ground truth data when working with private and personal information about individuals. Firstly, the risk of identifying individuals is reduced considerably, which enables the sharing of data for analysis amongst more organisations. Secondly, the fine tuning of synthetic datapoints to suit particular modelling and analyses could help to build more suitable models that can avoid biases found in the original ground truth data. In this paper we explore how a probabilistic synthetic data generator can be used to model data with high enough fidelity that it can be used to develop and validate state-of-the-art machine learning models. In particular, we use a Bayesian network model trained on gestational diabetes data, generated from a mobile health app collected from a number of health trusts in the UK. These data are used to train and test an established machine learning model developed by Sensyne Health using real-world data, and the resulting performance is compared to performance on ground truth data. In addition, a clinical validation is undertaken to explore if human experts can differentiate real patients from synthetic ones. We demonstrate that the Bayesian network synthetic data generator is able to mimic the ground truth closely enough to make it difficult for a human expert to distinguish between the two. We show that the data generator captures the interactions between features and the multivariate distributions close enough to enable classifiers to be inferred that imitate the key performance characteristics of models inferred from ground truth data. What is more, we demonstrate that the discovered mis-classifications found when testing using the synthetic data, are as informative as when testing using ground truth data.
Zhenchen Wang, Puja Myles, Anu Jain, James L. Keidel, Roberto Liddi, Lucy Mackillop, Carmelo Velardo, Allan Tucker
CBMS2
2021 Generating and evaluating cross-sectional synthetic electronic healthcare data: Preserving data utility and patient privacy
abstract
Abstract Electronic healthcare record data have been used to study risk factors of disease, treatment effectiveness and safety, and to inform healthcare service planning. There has been increasing interest in utilizing these data for new purposes such as for machine learning to develop predictive algorithms to aid diagnostic and treatment decisions. Synthetic data could potentially be an alternative to real‐world data for these purposes as well as reveal any biases in the data used for algorithm development. This article discusses the key requirements of synthetic data for multiple purposes and proposes an approach to generate and evaluate synthetic data focused on, but not limited to, cross‐sectional healthcare data. To our knowledge, this is the first article to propose a framework to generate and evaluate synthetic healthcare data with the aim of simultaneously preserving the complexities of ground truth data in the synthetic data while also ensuring privacy. We include findings and new insights from synthetic datasets modeled on both the Indian liver patient dataset and UK primary care dataset to demonstrate the application of this framework under different scenarios.
Zhenchen Wang, Puja Myles, Allan Tucker
Comput. Intell.2
2019 Generating and Evaluating Synthetic UK Primary Care Data: Preserving Data Utility & Patient Privacy
abstract
There is increasing interest in the potential of synthetic data to validate and benchmark machine learning algorithms as well as reveal any biases in real-world data used for algorithm development. This paper discusses the key requirements of synthetic data for such purposes and proposes an approach to generating and evaluating synthetic data that meets these requirements. We propose a framework to generate and evaluate synthetic data with the aim of simultaneously preserving the complexities of ground truth data in the synthetic data whilst also ensuring privacy. We include as a case study, a proof-of-concept synthetic dataset modelled on UK primary care data to demonstrate the application of this framework.
Zhenchen Wang, Puja Myles, Allan Tucker
CBMS2