Zhiyu Wan

dblp:169/1803 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
10since 2021 · last 2024
0000-0003-3752-5778ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Evaluating Fairness of Mask R-CNN for Kidney Infection Detection based on Renal Scintigraphy
abstract
99mTc-DMSA renal scan plays a crucial role in assessing functional abnormalities in the kidneys. A deep learning model, Mask R-CNN, showed much promise in diagnosing acute pyelonephritis, a type of kidney infection. This study evaluated the diagnostic performance and fairness of Mask R-CNN and Faster R-CNN using a99mTc-DMSA renal dataset. The classification results showed that Mask R-CNN achieved an accuracy of 0.89, while Faster R-CNN reached an accuracy of 0.88. Both models demonstrated strong classification capabilities for kidney conditions. Furthermore, the analysis of fairness across sex and age groups indicated that neither model exhibited significant bias, thereby supporting their suitability for clinical applications. Future research should consider integrating more patient data to further enhance the diagnostic capabilities and fairness assessments of the models.
Mingyan Wu, Ha Wu, Zhiyu Wan
IEEE Big Data5
2023 Managing re-identification risks while providing access to the All of Us research program
abstract
OBJECTIVE: The All of Us Research Program makes individual-level data available to researchers while protecting the participants' privacy. This article describes the protections embedded in the multistep access process, with a particular focus on how the data was transformed to meet generally accepted re-identification risk levels. METHODS: At the time of the study, the resource consisted of 329 084 participants. Systematic amendments were applied to the data to mitigate re-identification risk (eg, generalization of geographic regions, suppression of public events, and randomization of dates). We computed the re-identification risk for each participant using a state-of-the-art adversarial model specifically assuming that it is known that someone is a participant in the program. We confirmed the expected risk is no greater than 0.09, a threshold that is consistent with guidelines from various US state and federal agencies. We further investigated how risk varied as a function of participant demographics. RESULTS: The results indicated that 95th percentile of the re-identification risk of all the participants is below current thresholds. At the same time, we observed that risk levels were higher for certain race, ethnic, and genders. CONCLUSIONS: While the re-identification risk was sufficiently low, this does not imply that the system is devoid of risk. Rather, All of Us uses a multipronged data protection strategy that includes strong authentication practices, active monitoring of data misuse, and penalization mechanisms for users who violate terms of service.
Weiyi Xia, Melissa A. Basford, Robert J. Carroll, Ellen Wright Clayton, Paul A. Harris, Murat Kantarcioglu, Yongtai Liu, Steve Nyemba, Yevgeniy Vorobeychik, Zhiyu Wan, Bradley A. Malin
J. Am. Medical Informatics Assoc.10
2023 Defending Against Membership Inference Attacks on Beacon Services
abstract
Large genomic datasets are created through numerous activities, including recreational genealogical investigations, biomedical research, and clinical care. At the same time, genomic data has become valuable for reuse beyond their initial point of collection, but privacy concerns often hinder access. Beacon services have emerged to broaden accessibility to such data. These services enable users to query for the presence of a particular minor allele in a dataset, and information helps care providers determine if genomic variation is spurious or has some known clinical indication. However, various studies have shown that this process can leak information regarding if individuals are members of the underlying dataset. There are various approaches to mitigate this vulnerability, but they are limited in that they (1) typically rely on heuristics to add noise to the Beacon responses; (2) offer probabilistic privacy guarantees only, neglecting data utility; and (3) assume a batch setting where all queries arrive at once. In this article, we present a novel algorithmic framework to ensure privacy in a Beacon service setting with a minimal number of query response flips. We represent this problem as one of combinatorial optimization in both the batch setting and the online setting (where queries arrive sequentially). We introduce principled algorithms with both privacy and, in some cases, worst-case utility guarantees. Moreover, through extensive experiments, we show that the proposed approaches significantly outperform the state of the art in terms of privacy and utility, using a dataset consisting of 800 individuals and 1.3 million single nucleotide variants.
Rajagopal Venkatesaramani, Zhiyu Wan, Bradley A. Malin, Yevgeniy Vorobeychik
ACM Trans. Priv. Secur.2
2022 Assessing Machine Learning Based Generators for Synthetic Electronic Health Records: A Benchmarking
Chao Yan 0004, Ziqi Zhang 0005, Zhiyu Wan, Justin Guinney, Sean D. Mooney, Bradley A. Malin
AMIA4
2022 Supporting COVID-19 Disparity Investigations with Dynamically Adjusting Case Reporting Policies
J. Thomas Brown, Zhiyu Wan, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin
AMIA2
2022 A Scalable Tool for Realistic Health Data Re-identification Risk Assessment
Weiyi Xia, Yongtai Liu, Zhiyu Wan, Yevgeniy Vorobeychik, Murat Kantarcioglu, Ellen Wright Clayton, Bradley A. Malin
AMIA3
2022 Privacy-Preserving Publishing of Individual-Level Pandemic Data Based on a Game Theoretic Model
abstract
Sharing individual-level pandemic data is essential for accelerating the understanding of a disease. For example, COVID-19 data have been widely collected to support public health surveillance and research. In the United States, these data need to be de-identified before being released to the public due to privacy concerns. However, current data publishing approaches for individual-level pandemic data, such as those adopted by the U.S. Centers for Disease Control and Prevention (CDC), have not flexed over time to account for the dynamic nature of infection rates. Thus, the policies generated by these strategies may either raise privacy risks or impair the data utility (or usability). To optimize the tradeoff between privacy risk and data utility, we introduce a game theoretic model that adaptively generates policies to publish individual-level COVID-19 data according to infection dynamics. We model the data publishing process as a two-player Stackelberg game between a data publisher and a data recipient and then search for the best strategy for the publisher. In this game, we consider 1) the average accuracy of predicting future case counts for all demographic groups, and 2) the mutual information between the original data and the released data. We use COVID-19 case data from Vanderbilt University Medical Center from March 2020 to December 2021 to demonstrate our model and evaluate its effectiveness. The experimental results show that our game theoretic model outperforms all baseline approaches, including those adopted by CDC, while maintaining low privacy risk.
Abinitha Gourabathina, Zhiyu Wan, J. Thomas Brown, Chao Yan 0004, Bradley A. Malin
BIBM2
2022 How Adversarial Assumptions Influence Re-identification Risk Measures: A COVID-19 Case Study
Xinmeng Zhang, Zhiyu Wan, Chao Yan 0004, J. Thomas Brown, Weiyi Xia, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin
PSD2
2022 Dynamically adjusting case reporting policy to maximize privacy and public health utility in the face of a pandemic
abstract
OBJECTIVE: Supporting public health research and the public's situational awareness during a pandemic requires continuous dissemination of infectious disease surveillance data. Legislation, such as the Health Insurance Portability and Accountability Act of 1996 and recent state-level regulations, permits sharing deidentified person-level data; however, current deidentification approaches are limited. Namely, they are inefficient, relying on retrospective disclosure risk assessments, and do not flex with changes in infection rates or population demographics over time. In this paper, we introduce a framework to dynamically adapt deidentification for near-real time sharing of person-level surveillance data. MATERIALS AND METHODS: The framework leverages a simulation mechanism, capable of application at any geographic level, to forecast the reidentification risk of sharing the data under a wide range of generalization policies. The estimates inform weekly, prospective policy selection to maintain the proportion of records corresponding to a group size less than 11 (PK11) at or below 0.1. Fixing the policy at the start of each week facilitates timely dataset updates and supports sharing granular date information. We use August 2020 through October 2021 case data from Johns Hopkins University and the Centers for Disease Control and Prevention to demonstrate the framework's effectiveness in maintaining the PK11 threshold of 0.01. RESULTS: When sharing COVID-19 county-level case data across all US counties, the framework's approach meets the threshold for 96.2% of daily data releases, while a policy based on current deidentification techniques meets the threshold for 32.3%. CONCLUSION: Periodically adapting the data publication policies preserves privacy while enhancing public health utility through timely updates and sharing epidemiologically critical features.
J. Thomas Brown, Chao Yan 0004, Weiyi Xia, Zhijun Yin, Zhiyu Wan, Aris Gkoulalas-Divanis, Murat Kantarcioglu, Bradley A. Malin
J. Am. Medical Informatics Assoc.5
2021 De-identifying Socioeconomic Data at the Census Tract Level for Medical Research Through Constraint-based Clustering
Yongtai Liu, Douglas Conway, Zhiyu Wan, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin
AMIA3
2019 Biomedical Research Cohort Membership Disclosure on Social Media
Yongtai Liu, Chao Yan 0004, Zhijun Yin, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Bradley A. Malin
AMIA4
2018 Detecting the Presence of an Individual in Phenotypic Summary Data
Yongtai Liu, Zhiyu Wan, Weiyi Xia, Murat Kantarcioglu, Yevgeniy Vorobeychik, Ellen Wright Clayton, Abel N. Kho, David Carrell, Bradley A. Malin
AMIA2
2018 It's all in the timing: calibrating temporal penalties for biomedical data sharing
abstract
Objective: Biomedical science is driven by datasets that are being accumulated at an unprecedented rate, with ever-growing volume and richness. There are various initiatives to make these datasets more widely available to recipients who sign Data Use Certificate agreements, whereby penalties are levied for violations. A particularly popular penalty is the temporary revocation, often for several months, of the recipient's data usage rights. This policy is based on the assumption that the value of biomedical research data depreciates significantly over time; however, no studies have been performed to substantiate this belief. This study investigates whether this assumption holds true and the data science policy implications. Methods: This study tests the hypothesis that the value of data for scientific investigators, in terms of the impact of the publications based on the data, decreases over time. The hypothesis is tested formally through a mixed linear effects model using approximately 1200 publications between 2007 and 2013 that used datasets from the Database of Genotypes and Phenotypes, a data-sharing initiative of the National Institutes of Health. Results: The analysis shows that the impact factors for publications based on Database of Genotypes and Phenotypes datasets depreciate in a statistically significant manner. However, we further discover that the depreciation rate is slow, only ∼10% per year, on average. Conclusion: The enduring value of data for subsequent studies implies that revoking usage for short periods of time may not sufficiently deter those who would violate Data Use Certificate agreements and that alternative penalty mechanisms may need to be invoked.
Weiyi Xia, Zhiyu Wan, Zhijun Yin, James Gaupp, Yongtai Liu, Ellen Wright Clayton, Murat Kantarcioglu, Yevgeniy Vorobeychik, Bradley A. Malin
J. Am. Medical Informatics Assoc.2
2017 An Open Source Tool for Game Theoretic Health Data De-Identification
Fabian Prasser, James Gaupp, Zhiyu Wan, Weiyi Xia, Yevgeniy Vorobeychik, Murat Kantarcioglu, Klaus A. Kuhn, Bradley A. Malin
AMIA3
2015 Process-Driven Data Privacy
abstract
The quantity of personal data gathered by service providers via our daily activities continues to grow at a rapid pace. The sharing, and the subsequent analysis of, such data can support a wide range of activities, but concerns around privacy often prompt an organization to transform the data to meet certain protection models (e.g., k-anonymity or ε-differential privacy). These models, however, are based on simplistic adversarial frameworks, which can lead to both under- and over-protection. For instance, such models often assume that an adversary attacks a protected record exactly once. We introduce a principled approach to explicitly model the attack process as a series of steps. Specifically, we engineer a factored Markov decision process (FMDP) to optimally plan an attack from the adversary's perspective and assess the privacy risk accordingly. The FMDP captures the uncertainty in the adversary's belief (e.g., the number of identified individuals that match the de-identified data) and enables the analysis of various real world deterrence mechanisms beyond a traditional protection model, such as a penalty for committing an attack. We present an algorithm to solve the FMDP and illustrate its efficiency by simulating an attack on publicly accessible U.S. census records against a real identified resource of over 500,000 individuals in a voter registry. Our results demonstrate that while traditional privacy models commonly expect an adversary to attack exactly once per record, an optimal attack in our model may involve exploiting none, one, or more individuals in the pool of candidates, depending on context.
Weiyi Xia, Murat Kantarcioglu, Zhiyu Wan, Raymond Heatherly, Yevgeniy Vorobeychik, Bradley A. Malin
CIKM3