VLDB 2026 Research / reviewers in the wild / expert
Pei-Yuan Tsai
dblp:286/3101
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0001-4030-8953ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Biomedical knowledge graph verification with multitask learning architectures
Chih-Ping Wei, Pei-Yuan Tsai, Jih-Jane Li |
J. Biomed. Informatics | 2 |
| 2023 | User-Driven Synthetic Dataset Generation With Quantifiable Differential PrivacyabstractRecently, releasing data to a third party for secondary analysis has become a trend of service computing. However, data owners are concerned that such a move may expose individuals’ records, which is in violation of regulations such as the European Union's General Data Protection Regulation. Differential privacy has been proposed as a possible solution to the aforementioned problem. The privacy budget$\varepsilon$in differential privacy is for theoretical interpretation, but in practice, its application in measuring the risk of data disclosure has not been well studied, especially with sampling-based synthetic datasets. Moreover, datasets released by data owners with quantifiable privacy levels and the explicit utility for these datasets have yet to be well developed. In this paper, we present an intuitive approach for defining the privacy level (i.e., data hit rate and$k$-level) and utility level (i.e., basic statistics and a series of data mining models), and the privacy budget$\varepsilon$is quantified for evaluating the risk and utility of private data. In addition, we propose two user-driven synthetic dataset hunting methods to generate a synthetic dataset with the specified privacy objective, enabling the data owner (e.g., the government and financial companies) to understand the possible privacy risk and thereby release datasets with confirmed privacy level. To the best of our knowledge, this is the first method that allows data providers to automatically generate synthetic datasets with a quantifiable privacy level for the service of open data. Bo-Chen Tai, Yao-Tung Tsou, Szu-Chuang Li, Yennun Huang, Pei-Yuan Tsai, Yu-Cheng Tsai |
IEEE Trans. Serv. Comput. | 5 |
| 2021 | (k, ε , δ)-Anonymization: privacy-preserving data release based on k-anonymity and differential privacy
Yao-Tung Tsou, Mansour Naser Alraja, Li-Sheng Chen, Yu-Hsiang Chang, Yung-Li Hu, Yennun Huang, Chia-Mu Yu, Pei-Yuan Tsai |
Serv. Oriented Comput. Appl. | 8 |