EDBT 2026 Demo / reviewers in the wild / expert
Zhengli Huang
dblp:46/3690
· DBLP profile ↗
3ranked-venue papers
3as first author
0since 2021 · last 2008
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection › privacy-preserving data analysis
privacy-preserving data mining |
0.1 | 2 | 2008 | OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data Mining · ICDE 2008 Deriving Private Information from Randomized Data · SIGMOD Conference 2005 |
Privacy and data protection › differential privacy › local differential privacy
randomized response |
0.1 | 1 | 2008 | OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data Mining · ICDE 2008 |
Privacy and data protection › inference attack
data reconstruction |
0.1 | 1 | 2005 | Deriving Private Information from Randomized Data · SIGMOD Conference 2005 |
Privacy and data protection › information leakage
privacy breach |
0.1 | 1 | 2005 | Deriving Private Information from Randomized Data · SIGMOD Conference 2005 |
Privacy and data protection
randomization |
0.1 | 1 | 2005 | Deriving Private Information from Randomized Data · SIGMOD Conference 2005 |
Data mining › multivariate data analysis
correlation analysis |
0.0 | 1 | 2005 | Deriving Private Information from Randomized Data · SIGMOD Conference 2005 |
Methods — techniques the papers use, named apart from their topics
evolutionary multi-objective optimization · 0.2estimation theory · 0.2principal component analysis · 0.1bayes estimation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2008 | OptRR: Optimizing Randomized Response Schemes for Privacy-Preserving Data MiningabstractThe randomized response (RR) technique is a promising technique to disguise private categorical data in privacy-preserving data mining (PPDM). Although a number of RR-based methods have been proposed for various data mining computations, no study has systematically compared them to find optimal RR schemes. The difficulty of comparison lies in the fact that to compare two PPDM schemes, one needs to consider two conflicting metrics: privacy and utility. An optimal scheme based on one metric is usually the worst based on the other metric. In this paper, we first describe a method to quantify privacy and utility. We formulate the quantification as estimate problems, and use estimate theories to derive quantification. We then use an evolutionary multi-objective optimization method to find optimal disguise matrices for the randomized response technique. The experimental results have shown that our scheme has a much better performance than the existing RR schemes. Zhengli Huang, Wenliang Du 0001 |
ICDE | 1 |
| 2007 | Searching for Better Randomized Response Schemes for Privacy-Preserving Data Mining
Zhengli Huang, Wenliang Du 0001, Zhouxuan Teng |
PKDD | 1 |
| 2005 | Deriving Private Information from Randomized DataabstractRandomization has emerged as a useful technique for data disguising in privacy-preserving data mining. Its privacy properties have been studied in a number of papers. Kargupta et al. challenged the randomization schemes, and they pointed out that randomization might not be able to preserve privacy. However, it is still unclear what factors cause such a security breach, how they affect the privacy preserving property of the randomization, and what kinds of data have higher risk of disclosing their private contents even though they are randomized.We believe that the key factor is the correlations among attributes. We propose two data reconstruction methods that are based on data correlations. One method uses the Principal Component Analysis (PCA) technique, and the other method uses the Bayes Estimate (BE) technique. We have conducted theoretical and experimental analysis on the relationship between data correlations and the amount of private information that can be disclosed based our proposed data reconstructions schemes. Our studies have shown that when the correlations are high, the original data can be reconstructed more accurately, i.e., more private information can be disclosed.To improve privacy, we propose a modified randomization scheme, in which we let the correlation of random noises "similar" to the original data. Our results have shown that the reconstruction accuracy of both PCA-based and BE-based schemes become worse as the similarity increases. Zhengli Huang, Wenliang Du 0001, Biao Chen 0001 |
SIGMOD Conference | 1 |