Szu-Chuang Li

dblp:185/5249 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
3since 2021 · last 2023
0000-0001-6927-9562ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 6 · 1 first-author · 3 since 2021Security and privacy · 5 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2023 User-Driven Synthetic Dataset Generation With Quantifiable Differential Privacy
abstract
Recently, releasing data to a third party for secondary analysis has become a trend of service computing. However, data owners are concerned that such a move may expose individuals’ records, which is in violation of regulations such as the European Union's General Data Protection Regulation. Differential privacy has been proposed as a possible solution to the aforementioned problem. The privacy budget$\varepsilon$in differential privacy is for theoretical interpretation, but in practice, its application in measuring the risk of data disclosure has not been well studied, especially with sampling-based synthetic datasets. Moreover, datasets released by data owners with quantifiable privacy levels and the explicit utility for these datasets have yet to be well developed. In this paper, we present an intuitive approach for defining the privacy level (i.e., data hit rate and$k$-level) and utility level (i.e., basic statistics and a series of data mining models), and the privacy budget$\varepsilon$is quantified for evaluating the risk and utility of private data. In addition, we propose two user-driven synthetic dataset hunting methods to generate a synthetic dataset with the specified privacy objective, enabling the data owner (e.g., the government and financial companies) to understand the possible privacy risk and thereby release datasets with confirmed privacy level. To the best of our knowledge, this is the first method that allows data providers to automatically generate synthetic datasets with a quantifiable privacy level for the service of open data.
Bo-Chen Tai, Yao-Tung Tsou, Szu-Chuang Li, Yennun Huang, Pei-Yuan Tsai, Yu-Cheng Tsai
IEEE Trans. Serv. Comput.3
2022 Examining the Utility of Differentially Private Synthetic Data Generated using Variational Autoencoder with TensorFlow Privacy
abstract
With the emergence of AI(artificial intelligence), it is becoming more and more critical for organizations to utilize it to their advantage. However, organizations that possess a decent amount of data might not have the technical competence to perform machine learning, and vice versa. Hence, it is reasonable for the two kinds of organizations to work together to realize the value of the data. With the increasing concern over data privacy, regulations such as GDPR(General Data Protection Regulation) prevent an organization from sharing data with another unless the data is processed to the point that the individuals in the data are not identifiable. Various ways of data anonymization have been proposed and developed, including the ones that utilize neural networks to achieve the goal, like AE, VAE, and GAN. With the addition of a differential privacy framework like TensorFlow Privacy, privacy can be guaranteed, but data still needs to be usable after privacy protection measures are deployed. The present study aims to integrate TensorFlow Privacy into the synthetic data generation process and evaluate its usefulness for daily use in the industries. Since TensorFlow Privacy brings a provable privacy guarantee to synthetic data, the present study focuses on the evaluation of data utility. TensorFlow is widely used for machine learning in the industry and academically. TensorFlow Privacy, which is also developed by Google, can prove to be a valuable addition to the synthetic data generation pipeline. The result shows that VAE with TensorFlow Privacy 1) generates synthetic data with good data utility in most cases in terms of descriptive statistics and machine learning classification tasks, and 2) The customizable TensorFlow Privacy parameters work as intended in terms of privacy-utility trade-off.
Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, Pang-Chieh Wang
PRDC2
2021 A VAE Conversion Method for Private Data Linkage
abstract
Data linkage plays a crucial role in realizing big data's value but is often regarded as a threat to personal privacy. Regulations like GDPR requires users' consent on each specific use of data, which is not practical for data analyzers. In this study, we propose a way to address the problem by having a trustworthy third party collect data from two or more parties, then use the data to train one or more variational autoencoder (VAE) models to remove privacy and send them to the data providers. Using this model, the users express their consent to share data with a trustworthy party. The third party links data from various datasets together to build a variational autoencoder model that allows all parties to generate datasets with full attributes without revealing sensitive personal data. System architectures and machine learning accuracy of generated data sets are measured in this study.
Bo-Chen Tai, Szu-Chuang Li, Yennun Huang
PRDC2
2019 Evaluating Variational Autoencoder as a Private Data Release Mechanism for Tabular Data
abstract
Multi-market businesses can collect data from different business entities and aggregate data from various sources to create value. However, due to the restriction of privacy regulation, it could be illegal to exchange data between business entities of the same parent company, unless the users have opted-in to allow it. Regulations such as the EU's GDPR allows data exchange if data is anonymized appropriately. In this study, we use variational autoencoder as a mechanism to generate synthetic data. The privacy and utility of the generated data sets are measured. And its performance is compared with the performance of the plain autoencoder. The primary findings of this study are 1) variational autoencoder can be an option for data exchange with good accuracy even when the number of latent dimensions is low 2) plain autoencoder still provides better accuracy when the number of hidden nodes is high 3) variational autoencoder, as a generative model, can be given to a data user to generate his version of data that closely mimic the original data set.
Szu-Chuang Li, Bo-Chen Tai, Yennun Huang
PRDC1
2018 Exploring the Relationship Between Dimensionality Reduction and Private Data Release
abstract
It is important to facilitate data sharing between data owners and data analysts as data owners do not always have the ability to process and analyze data. For example, governments around the world are starting to release collected data to the public to leverage data analysis competence of the crowd. However, some privacy leakage incidents have made the public to rediscover the importance of privacy protection, leading to new privacy regulations. In existing researches dimensionality reduction has played an important role in private data release mechanisms to improve utility but its influence on privacy protection has never been examined. In this study, we perform a series of experiments and found that dimensionality reduction could provide similar privacy protection effects as K-anonymity mechanisms, and it could work as a preprocessor of K-anonymity process to it to reduce the generalization and suppression needed.
Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, Neeraj Suri, Pang-Chieh Wang
PRDC2
2017 Data-Driven Approach for Evaluating Risk of Disclosure and Utility in Differentially Private Data Release
abstract
Differential privacy (DP) is a popular technique for protecting individual privacy and at the same for releasing data for public use. However, very few research efforts are devoted to the balance between the corresponding risk of data disclosure (RoD) and data utility. In this paper, we propose data-driven approaches for differentially private data release to evaluate RoD, and offer algorithms to evaluate whether the differentially private synthetic dataset has sufficient privacy. In addition to the privacy, the utility of the synthetic dataset is an important metric for differentially private data release. Thus, we also propose the data-driven algorithm via curve fitting to measure and predict the error of the statistical result incurred by random noise added to the original dataset. Finally, we present an algorithm for choosing appropriate privacy budget ∈ with the balance between the privacy and utility.
Kang-Cheng Chen, Chia-Mu Yu, Bo-Chen Tai, Szu-Chuang Li, Yao-Tung Tsou, Yennun Huang, Chia-Ming Lin
AINA4
2017 K-Aggregation: Improving Accuracy for Differential Privacy Synthetic Dataset by Utilizing K-Anonymity Algorithm
abstract
Enterprises and governments around the world have been attempting to leverage intelligence from the community by making formally in-house database available to the public for analyzing. The released data was often “anonymized”: sensitive attributes were removed from the dataset for privacy protection. However it is proved that masking sensitive attributes alone is not adequate for data protection. Differential privacy can be used to generate “synthetic dataset” that retain statistical properties of the original dataset and limit data-leaking risk at the same time, but there's always a trade-off between data privacy and utility. In this study we aggregate data counts across value with little counts to ease the problem of excessive error at the data value with small data count. Experiments show that K-aggregation has the potential to reduce error of count queries on value with smaller counts. Limitations of this approach are also discussed.
Bo-Chen Tai, Szu-Chuang Li, Yennun Huang
AINA2
2017 Evaluating the Risk of Data Disclosure Using Noise Estimation for Differential Privacy
abstract
Differential privacy is a recent notion of data privacy protection, which does not matter even when an attacker has arbitrary background knowledge in advance. Consequently, it is viewed as a reliable protection mechanism for sensitive information. Differential privacy introduces Laplace noise to hide the true value in a dataset while preserving statistic properties. However, the large amount of Laplace noise added into a dataset is typically defined by the discursive scale parameter of the Laplace distribution. The privacy parameter ε in differential privacy is with theoretical interpretation, but the implication on the risk of data disclosure (called RoD for short) in practice has not yet been studied. Moreover, choosing appropriate value for ε is not an easy task since it impacts the level of privacy in a dataset significantly. In this paper, we define and evaluate the RoD in a dataset with either numerical or binary attributes for numerical or counting queries with multiple attributes based on the noise estimation. Through confidence probability of noise estimation, we give a simple way to choose the privacy parameter ε. Finally, we show the relation of the RoD and privacy parameter ε in experimental results. To the best of our knowledge, this is the first research work in using noise estimation to practically evaluate the RoD for multiple attributes (both numerical and binary data).
Hung-Li Chen, Jia-Yang Chen, Yao-Tung Tsou, Chia-Mu Yu, Bo-Chen Tai, Szu-Chuang Li, Yennun Huang, Chia-Ming Lin
PRDC6