Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhengli Huang

dblp:46/3690 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
0since 2021 · last 2008
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
2 papers
Privacy and data protection · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection › privacy-preserving data analysis
privacy-preserving data mining
0.122008
OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data Mining · ICDE 2008
Deriving Private Information from Randomized Data · SIGMOD Conference 2005
Privacy and data protection › differential privacy › local differential privacy
randomized response
0.112008
OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data Mining · ICDE 2008
Privacy and data protection › inference attack
data reconstruction
0.112005
Deriving Private Information from Randomized Data · SIGMOD Conference 2005
Privacy and data protection › information leakage
privacy breach
0.112005
Deriving Private Information from Randomized Data · SIGMOD Conference 2005
Privacy and data protection
randomization
0.112005
Deriving Private Information from Randomized Data · SIGMOD Conference 2005
Data mining › multivariate data analysis
correlation analysis
0.012005
Deriving Private Information from Randomized Data · SIGMOD Conference 2005

Methods — techniques the papers use, named apart from their topics

evolutionary multi-objective optimization · 0.2estimation theory · 0.2principal component analysis · 0.1bayes estimation · 0.1
YearPublicationVenuePosition
2008 OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data Mining
abstract
The randomized response (RR) technique is a promising technique to disguise private categorical data in privacy-preserving data mining (PPDM). Although a number of RR-based methods have been proposed for various data mining computations, no study has systematically compared them to find optimal RR schemes. The difficulty of comparison lies in the fact that to compare two PPDM schemes, one needs to consider two conflicting metrics: privacy and utility. An optimal scheme based on one metric is usually the worst based on the other metric. In this paper, we first describe a method to quantify privacy and utility. We formulate the quantification as estimate problems, and use estimate theories to derive quantification. We then use an evolutionary multi-objective optimization method to find optimal disguise matrices for the randomized response technique. The experimental results have shown that our scheme has a much better performance than the existing RR schemes.
Zhengli Huang, Wenliang Du 0001
ICDE1
2007 Searching for Better Randomized Response Schemes for Privacy-Preserving Data Mining
Zhengli Huang, Wenliang Du 0001, Zhouxuan Teng
PKDD1
2005 Deriving Private Information from Randomized Data
abstract
Randomization has emerged as a useful technique for data disguising in privacy-preserving data mining. Its privacy properties have been studied in a number of papers. Kargupta et al. challenged the randomization schemes, and they pointed out that randomization might not be able to preserve privacy. However, it is still unclear what factors cause such a security breach, how they affect the privacy preserving property of the randomization, and what kinds of data have higher risk of disclosing their private contents even though they are randomized.We believe that the key factor is the correlations among attributes. We propose two data reconstruction methods that are based on data correlations. One method uses the Principal Component Analysis (PCA) technique, and the other method uses the Bayes Estimate (BE) technique. We have conducted theoretical and experimental analysis on the relationship between data correlations and the amount of private information that can be disclosed based our proposed data reconstructions schemes. Our studies have shown that when the correlations are high, the original data can be reconstructed more accurately, i.e., more private information can be disclosed.To improve privacy, we propose a modified randomization scheme, in which we let the correlation of random noises "similar" to the original data. Our results have shown that the reconstruction accuracy of both PCA-based and BE-based schemes become worse as the similarity increases.
Zhengli Huang, Wenliang Du 0001, Biao Chen 0001
SIGMOD Conference1