Chao Li 0003

dblp:66/190-3 · DBLP profile ↗
← Back
12ranked-venue papers
9as first author
0since 2021 · last 2015
0000-0002-9578-4316ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 8 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
7 papers
Privacy and data protection · 100%
Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 84% Data mining · 10% Web and social media mining · 6%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
differential privacy
0.962015
The matrix mechanism: optimizing linear counting queries under differential privacy · VLDB J. 2015
A Theory of Pricing Private Data · ACM Trans. Database Syst. 2014
A Data- and Workload-Aware Query Answering Algorithm for Range Queries Under Differential Privacy · Proc. VLDB Endow. 2014
Query processing and optimization
range query
0.212014
A Data- and Workload-Aware Query Answering Algorithm for Range Queries Under Differential Privacy · Proc. VLDB Endow. 2014
Privacy and data protection › data sharing
data market
0.212014
A Theory of Pricing Private Data · ACM Trans. Database Syst. 2014
Privacy and data protection › differential privacy
differentially private query answering
0.212014
A Data- and Workload-Aware Query Answering Algorithm for Range Queries Under Differential Privacy · Proc. VLDB Endow. 2014
Privacy and data protection › differential privacy › differentially private query answering
counting queries
0.112012
An Adaptive Mechanism for Accurate Query Answering under Differential Privacy · Proc. VLDB Endow. 2012
Query processing and optimization › secure query processing
differentially private query answering
0.112010
Optimizing linear counting queries under differential privacy · PODS 2010
Query processing and optimization › multi-query optimization
query workload optimization
0.112010
Optimizing linear counting queries under differential privacy · PODS 2010
Privacy and data protection
anonymization
0.112010
Resisting structural re-identification in anonymized social networks · VLDB J. 2010
Web and social media mining › social network analysis
social network
0.012010
Resisting structural re-identification in anonymized social networks · VLDB J. 2010
Data mining › structured data mining
graph mining
0.012009
Accurate Estimation of the Degree Distribution of Private Networks · ICDM 2009
Data mining
network analysis
0.012009
Accurate Estimation of the Degree Distribution of Private Networks · ICDM 2009

Methods — techniques the papers use, named apart from their topics

differential privacy · 0.6matrix mechanism · 0.4mechanism design · 0.4data-dependent partitioning · 0.4bucket count estimation · 0.4strategy query selection · 0.1
YearPublicationVenuePosition
2015 Lower Bounds on the Error of Query Sets Under the Differentially-Private Matrix Mechanism
Chao Li 0003, Gerome Miklau
Theory Comput. Syst.1
2015 The matrix mechanism: optimizing linear counting queries under differential privacy
Chao Li 0003, Gerome Miklau, Michael Hay, Andrew McGregor 0001, Vibhor Rastogi
VLDB J.1
2014 A Data- and Workload-Aware Query Answering Algorithm for Range Queries Under Differential Privacy
abstract
We describe a new algorithm for answering a given set of range queries under ε-differential privacy which often achieves substantially lower error than competing methods. Our algorithm satisfies differential privacy by adding noise that is adapted to the input data and to the given query set. We first privately learn a partitioning of the domain into buckets that suit the input data well. Then we privately estimate counts for each bucket, doing so in a manner well-suited for the given query set. Since the performance of the algorithm depends on the input database, we evaluate it on a wide range of real datasets, showing that we can achieve the benefits of data-dependence on both "easy" and "hard" databases.
Chao Li 0003, Michael Hay, Gerome Miklau, Yue Wang 0070
Proc. VLDB Endow.1
2014 A Theory of Pricing Private Data
abstract
Personal data has value to both its owner and to institutions who would like to analyze it. Privacy mechanisms protect the owner's data while releasing to analysts noisy versions of aggregate query results. But such strict protections of the individual's data have not yet found wide use in practice. Instead, Internet companies, for example, commonly provide free services in return for valuable sensitive information from users, which they exploit and sometimes sell to third parties. As awareness of the value of personal data increases, so has the drive to compensate the end-user for her private information. The idea of monetizing private data can improve over the narrower view of hiding private data, since it empowers individuals to control their data through financial means. In this article we propose a theoretical framework for assigning prices to noisy query answers as a function of their accuracy, and for dividing the price amongst data owners who deserve compensation for their loss of privacy. Our framework adopts and extends key principles from both differential privacy and query pricing in data markets. We identify essential properties of the pricing function and micropayments, and characterize valid solutions.
Chao Li 0003, Daniel Yang Li, Gerome Miklau, Dan Suciu
ACM Trans. Database Syst.1
2013 A theory of pricing private data
abstract
Personal data has value to both its owner and to institutions who would like to analyze it. Privacy mechanisms protect the owner's data while releasing to analysts noisy versions of aggregate query results. But such strict protections of individual's data have not yet found wide use in practice. Instead, Internet companies, for example, commonly provide free services in return for valuable sensitive information from users, which they exploit and sometimes sell to third parties.
Chao Li 0003, Daniel Yang Li, Gerome Miklau, Dan Suciu
ICDT1
2013 Optimal error of query sets under the differentially-private matrix mechanism
abstract
A common goal of privacy research is to release synthetic data that satisfies a formal privacy guarantee and can be used by an analyst in place of the original data. To achieve reasonable accuracy, a synthetic data set must be tuned to support a specified set of queries accurately, sacrificing fidelity for other queries.
Chao Li 0003, Gerome Miklau
ICDT1
2012 Pricing Aggregate Queries in a Data Marketplace
Chao Li 0003, Gerome Miklau
WebDB1
2012 An Adaptive Mechanism for Accurate Query Answering under Differential Privacy
abstract
We propose a novel mechanism for answering sets of counting queries under differential privacy. Given a workload of counting queries, the mechanism automatically selects a different set of "strategy" queries to answer privately, using those answers to derive answers to the workload. The main algorithm proposed in this paper approximates the optimal strategy for any workload of linear counting queries. With no cost to the privacy guarantee, the mechanism improves significantly on prior approaches and achieves near-optimal error for many workloads, when applied under (ε, δ)-differential privacy. The result is an adaptive mechanism which can help users achieve good utility without requiring that they reason carefully about the best formulation of their task.
Chao Li 0003, Gerome Miklau
Proc. VLDB Endow.1
2010 Automated Tuning in Parallel Sorting on Multi-core Architectures
Chao Li 0003, Ninghe Pan, Xiaotong Zhuang, Ling Shao 0002
Euro-Par (1)2
2010 Optimizing linear counting queries under differential privacy
abstract
Differential privacy is a robust privacy standard that has been successfully applied to a range of data analysis tasks. But despite much recent work, optimal strategies for answering a collection of related queries are not known.
Chao Li 0003, Michael Hay, Vibhor Rastogi, Gerome Miklau, Andrew McGregor 0001
PODS1
2010 Resisting structural re-identification in anonymized social networks
Michael Hay, Gerome Miklau, David D. Jensen, Don Towsley, Chao Li 0003
VLDB J.5
2009 Accurate Estimation of the Degree Distribution of Private Networks
abstract
We describe an efficient algorithm for releasing a provably private estimate of the degree distribution of a network. The algorithm satisfies a rigorous property of differential privacy, and is also extremely efficient, running on networks of 100 million nodes in a few seconds. Theoretical analysis shows that the error scales linearly with the number of unique degrees, whereas the error of conventional techniques scales linearly with the number of nodes. We complement the theoretical analysis with a thorough empirical analysis on real and synthetic graphs, showing that the algorithm's variance and bias is low, that the error diminishes as the size of the input graph increases, and that common analyses like fitting a power-law can be carried out very accurately.
Michael Hay, Chao Li 0003, Gerome Miklau, David D. Jensen
ICDM2