Ziqi Zhang 0005

dblp:97/7236-5 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
8since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 8 since 2021
YearPublicationVenuePosition
2022 Assessing Machine Learning Based Generators for Synthetic Electronic Health Records: A Benchmarking
Chao Yan 0004, Ziqi Zhang 0005, Zhiyu Wan, Justin Guinney, Sean D. Mooney, Bradley A. Malin
AMIA3
2022 Keeping synthetic patients on track: feedback mechanisms to mitigate performance drift in longitudinal health data simulation
abstract
OBJECTIVE: Synthetic data are increasingly relied upon to share electronic health record (EHR) data while maintaining patient privacy. Current simulation methods can generate longitudinal data, but the results are unreliable for several reasons. First, the synthetic data drifts from the real data distribution over time. Second, the typical approach to quality assessment, which is based on the extent to which real records can be distinguished from synthetic records using a critic model, often fails to recognize poor simulation results. In this article, we introduce a longitudinal simulation framework, called LS-EHR, which addresses these issues. MATERIALS AND METHODS: LS-EHR enhances simulation through conditional fuzzing and regularization, rejection sampling, and prior knowledge embedding. We compare LS-EHR to the state-of-the-art using data from 60 000 EHRs from Vanderbilt University Medical Center (VUMC) and the All of Us Research Program. We assess discrimination between real and synthetic data over time. We evaluate the generation process and critic model using the area under the receiver operating characteristic curve (AUROC). For the critic, a higher value indicates a more robust model for quality assessment. For the generation process, a lower value indicates better synthetic data quality. RESULTS: The LS-EHR critic improves discrimination AUROC from 0.655 to 0.909 and 0.692 to 0.918 for VUMC and All of Us data, respectively. By using the new critic, the LS-EHR generation model reduces the AUROC from 0.909 to 0.758 and 0.918 to 0.806. CONCLUSION: LS-EHR can substantially improve the usability of simulated longitudinal EHR data.
Ziqi Zhang 0005, Chao Yan 0004, Bradley A. Malin
J. Am. Medical Informatics Assoc.1
2022 Forecasting the future clinical events of a patient through contrastive learning
abstract
OBJECTIVE: Deep learning models for clinical event forecasting (CEF) based on a patient's medical history have improved significantly over the past decade. However, their transition into practice has been limited, particularly for diseases with very low prevalence. In this paper, we introduce CEF-CL, a novel method based on contrastive learning to forecast in the face of a limited number of positive training instances. MATERIALS AND METHODS: CEF-CL consists of two primary components: (1) unsupervised contrastive learning for patient representation and (2) supervised transfer learning over the derived representation. We evaluate the new method along with state-of-the-art model architectures trained in a supervised manner with electronic health records data from Vanderbilt University Medical Center and the All of Us Research Program, covering 48 000 and 16 000 patients, respectively. We assess forecasting for over 100 diagnosis codes with respect to their area under the receiver operator characteristic curve (AUROC) and area under the precision-recall curve (AUPRC). We investigate the correlation between forecasting performance improvement and code prevalence via a Wald Test. RESULTS: CEF-CL achieved an average AUROC and AUPRC performance improvement over the state-of-the-art of 8.0%-9.3% and 11.7%-32.0%, respectively. The improvement in AUROC was negatively correlated with the number of positive training instances (P < .001). CONCLUSION: This investigation indicates that clinical event forecasting can be improved significantly through contrastive representation learning, especially when the number of positive training instances is small.
Ziqi Zhang 0005, Chao Yan 0004, Xinmeng Zhang, Steve Nyemba, Bradley A. Malin
J. Am. Medical Informatics Assoc.1
2022 Membership inference attacks against synthetic health data
Ziqi Zhang 0005, Chao Yan 0004, Bradley A. Malin
J. Biomed. Informatics1
2021 Synthetic Data to Support Engineering and Demonstrations in the All of Us Research Program
Chao Yan 0004, Steve Nyemba, Kelsey R. Mayo, Ziqi Zhang 0005, Francis Ratsimbazafy, Bradley A. Malin
AMIA4
2021 CCF-CL: Forecasting the Clinical Status of a Patient Through Contrastive Learning
Ziqi Zhang 0005, Chao Yan 0004, Xinmeng Zhang, Steve Nyemba, Bradley A. Malin
AMIA1
2021 Predicting brain function status changes in critically ill patients via Machine learning
abstract
OBJECTIVE: In intensive care units (ICUs), a patient's brain function status can shift from a state of acute brain dysfunction (ABD) to one that is ABD-free and vice versa, which is challenging to forecast and, in turn, hampers the allocation of hospital resources. We aim to develop a machine learning model to predict next-day brain function status changes. MATERIALS AND METHODS: Using multicenter prospective adult cohorts involving medical and surgical ICU patients from 2 civilian and 3 Veteran Affairs hospitals, we trained and externally validated a light gradient boosting machine to predict brain function status changes. We compared the performances of the boosting model against state-of-the-art models-an ABD predictive model and its variants. We applied Shapley additive explanations to identify influential factors to develop a compact model. RESULTS: There were 1026 critically ill patients without evidence of prior major dementia, or structural brain diseases, from whom 12 295 daily transitions (ABD: 5847 days; ABD-free: 6448 days) were observed. The boosting model achieved an area under the receiver-operating characteristic curve (AUROC) of 0.824 (95% confidence interval [CI], 0.821-0.827), compared with the state-of-the-art models of 0.697 (95% CI, 0.693-0.701) with P < .001. Using 13 identified top influential factors, the compact model achieved 99.4% of the boosting model on AUROC. The boosting and the compact models demonstrated high generalizability in external validation by achieving an AUROC of 0.812 (95% CI, 0.812-0.813). CONCLUSION: The inputs of the compact model are based on several simple questions that clinicians can quickly answer in practice, which demonstrates the model has direct prospective deployment potential into clinical practice, aiding in critical hospital resource allocation.
Chao Yan 0004, Ziqi Zhang 0005, Wencong Chen, Bradley A. Malin, Eugene Wesley Ely, Mayur B. Patel, You Chen 0001
J. Am. Medical Informatics Assoc.3
2021 SynTEG: a framework for temporal structured electronic health data simulation
abstract
OBJECTIVE: Simulating electronic health record data offers an opportunity to resolve the tension between data sharing and patient privacy. Recent techniques based on generative adversarial networks have shown promise but neglect the temporal aspect of healthcare. We introduce a generative framework for simulating the trajectory of patients' diagnoses and measures to evaluate utility and privacy. MATERIALS AND METHODS: The framework simulates date-stamped diagnosis sequences based on a 2-stage process that 1) sequentially extracts temporal patterns from clinical visits and 2) generates synthetic data conditioned on the learned patterns. We designed 3 utility measures to characterize the extent to which the framework maintains feature correlations and temporal patterns in clinical events. We evaluated the framework with billing codes, represented as phenome-wide association study codes (phecodes), from over 500 000 Vanderbilt University Medical Center electronic health records. We further assessed the privacy risks based on membership inference and attribute disclosure attacks. RESULTS: The simulated temporal sequences exhibited similar characteristics to real sequences on the utility measures. Notably, diagnosis prediction models based on real versus synthetic temporal data exhibited an average relative difference in area under the ROC curve of 1.6% with standard deviation of 3.8% for 1276 phecodes. Additionally, the relative difference in the mean occurrence age and time between visits were 4.9% and 4.2%, respectively. The privacy risks in synthetic data, with respect to the membership and attribute inference were negligible. CONCLUSION: This investigation indicates that temporal diagnosis code sequences can be simulated in a manner that provides utility and respects privacy.
Ziqi Zhang 0005, Chao Yan 0004, Thomas A. Lasko, Jimeng Sun 0001, Bradley A. Malin
J. Am. Medical Informatics Assoc.1
2020 Generating Electronic Health Records with Multiple Data Types and Constraints
Chao Yan 0004, Ziqi Zhang 0005, Steve Nyemba, Bradley A. Malin
AMIA2
2020 Ensuring electronic medical record simulation through better training, modeling, and evaluation
abstract
OBJECTIVE: Electronic medical records (EMRs) can support medical research and discovery, but privacy risks limit the sharing of such data on a wide scale. Various approaches have been developed to mitigate risk, including record simulation via generative adversarial networks (GANs). While showing promise in certain application domains, GANs lack a principled approach for EMR data that induces subpar simulation. In this article, we improve EMR simulation through a novel pipeline that (1) enhances the learning model, (2) incorporates evaluation criteria for data utility that informs learning, and (3) refines the training process. MATERIALS AND METHODS: We propose a new electronic health record generator using a GAN with a Wasserstein divergence and layer normalization techniques. We designed 2 utility measures to characterize similarity in the structural properties of real and simulated EMRs in the original and latent space, respectively. We applied a filtering strategy to enhance GAN training for low-prevalence clinical concepts. We evaluated the new and existing GANs with utility and privacy measures (membership and disclosure attacks) using billing codes from over 1 million EMRs at Vanderbilt University Medical Center. RESULTS: The proposed model outperformed the state-of-the-art approaches with significant improvement in retaining the nature of real records, including prediction performance and structural properties, without sacrificing privacy. Additionally, the filtering strategy achieved higher utility when the EMR training dataset was small. CONCLUSIONS: These findings illustrate that EMR simulation through GANs can be substantially improved through more appropriate training, modeling, and evaluation criteria.
Ziqi Zhang 0005, Chao Yan 0004, Diego Mesa, Jimeng Sun 0001, Bradley A. Malin
J. Am. Medical Informatics Assoc.1