Cheryl J. Flynn

dblp:192/5187 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
2since 2021 · last 2023
0009-0001-2069-8907ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSecurity and privacy · 1Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
4 papers
Privacy and data protection · 93% Cryptographic protocols and secure computation · 7%
Databases, data mining, and information retrieval
3 papers
Graph data management · 48% Data mining · 32% Distributed and cloud data management · 21%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
differential privacy
2.042023
Local dampening: differential privacy for non-numeric queries via local sensitivity · VLDB J. 2023
Global and Local Differentially Private Release of Count-Weighted Graphs · Proc. ACM Manag. Data 2023
Local Dampening: Differential Privacy for Non-numeric Queries via Local Sensitivity · Proc. VLDB Endow. 2020
Privacy and data protection › differential privacy › privacy mechanism design › sensitivity analysis
local sensitivity
1.122023
Local dampening: differential privacy for non-numeric queries via local sensitivity · VLDB J. 2023
Local Dampening: Differential Privacy for Non-numeric Queries via Local Sensitivity · Proc. VLDB Endow. 2020
Privacy and data protection › differential privacy
local differential privacy
0.712023
Global and Local Differentially Private Release of Count-Weighted Graphs · Proc. ACM Manag. Data 2023
Privacy and data protection
privacy-preserving record linkage
0.312017
Composing Differential Privacy and Secure Computation: A Case Study on Scaling Private Record Linkage · CCS 2017
Cryptographic protocols and secure computation › secure multiparty computation
secure two-party computation
0.312017
Composing Differential Privacy and Secure Computation: A Case Study on Scaling Private Record Linkage · CCS 2017
Graph data management
graph release
0.212023
Global and Local Differentially Private Release of Count-Weighted Graphs · Proc. ACM Manag. Data 2023

Methods — techniques the papers use, named apart from their topics

post-processing · 1.3local sensitivity · 0.9exponential mechanism · 0.9secure two-party computation · 0.6differential privacy · 0.6
YearPublicationVenuePosition
2023 Global and Local Differentially Private Release of Count-Weighted Graphs
abstract
Many complex natural and technological systems are commonly modeled as count-weighted graphs, where nodes represent entities, edges model relationships between them, and edge weights define some counting statistics associated with each relationship. As graph data usually contain sensitive information about entities, preserving privacy when releasing this type of data becomes an important issue. In this context, differential privacy (DP) has become the de facto standard for data release under strong privacy guarantees. When dealing with DP for weighted graphs, most state-of-the-art works assume that the graph topology is known. However, in several real-world applications, the privacy of the graph topology also needs to be ensured. In this paper, we aim to bridge the gap between DP and count-weighted graph data release, considering both graph structure and edge weights as private information. We first adapt the weighted graph DP definition to take into account the privacy of the graph structure. We then develop two novel approaches to privately releasing count-weighted graphs under the notions of global and local DP. We also leverage the post-processing property of DP to improve the accuracy of the proposed techniques considering graph domain constraints. Experiments using real-world graph data demonstrate the superiority of our approaches in terms of utility over existing techniques, enabling subsequent computation of a variety of statistics on the released graph with high utility, in some cases comparable to the non-private results.
Felipe T. Brito, Victor A. E. de Farias, Cheryl J. Flynn, Subhabrata Majumdar, Javam C. Machado, Divesh Srivastava
Proc. ACM Manag. Data3
2023 Local dampening: differential privacy for non-numeric queries via local sensitivity
Victor A. E. de Farias, Felipe T. Brito, Cheryl J. Flynn, Javam C. Machado, Subhabrata Majumdar, Divesh Srivastava
VLDB J.3
2020 Local Dampening: Differential Privacy for Non-numeric Queries via Local Sensitivity
abstract
Differential privacy is the state-of-the-art formal definition for data release under strong privacy guarantees. A variety of mechanisms have been proposed in the literature for releasing the noisy output of numeric queries (e.g., using the Laplace mechanism), based on the notions of global sensitivity and local sensitivity. However, although there has been some work on generic mechanisms for releasing the output of non-numeric queries using global sensitivity (e.g., the Exponential mechanism), the literature lacks generic mechanisms for releasing the output of non-numeric queries using local sensitivity to reduce the noise in the query output. In this work, we remedy this shortcoming and present the local dampening mechanism. We adapt the notion of local sensitivity for the non-numeric setting and leverage it to design a generic non-numeric mechanism. We illustrate the effectiveness of the local dampening mechanism by applying it to two diverse problems: (i) Influential node analysis. Given an influence metric, we release the top-k most influential nodes while preserving the privacy of the relationship between nodes in the network; (ii) Decision tree induction. We provide a private adaptation to the ID3 algorithm to build decision trees from a given tabular dataset. Experimental results show that we could reduce the use of privacy budget by 3 to 4 orders of magnitude for Influential node analysis and increase accuracy up to 12% for Decision tree induction when compared to global sensitivity based approaches.
Victor A. E. de Farias, Felipe T. Brito, Cheryl J. Flynn, Javam C. Machado, Subhabrata Majumdar, Divesh Srivastava
Proc. VLDB Endow.3
2017 Composing Differential Privacy and Secure Computation: A Case Study on Scaling Private Record Linkage
abstract
Private record linkage (PRL) is the problem of identifying pairs of records that are similar as per an input matching rule from databases held by two parties that do not trust one another. We identify three key desiderata that a PRL solution must ensure: (1) perfect precision and high recall of matching pairs, (2) a proof of end-to-end privacy, and (3) communication and computational costs that scale subquadratically in the number of input records. We show that all of the existing solutions for PRL? including secure 2-party computation (S2PC), and their variants that use non-private or differentially private (DP) blocking to ensure subquadratic cost -- violate at least one of the three desiderata. In particular, S2PC techniques guarantee end-to-end privacy but have either low recall or quadratic cost. In contrast, no end-to-end privacy guarantee has been formalized for solutions that achieve subquadratic cost. This is true even for solutions that compose DP and S2PC: DP does not permit the release of any exact information about the databases, while S2PC algorithms for PRL allow the release of matching records.
Xi He 0001, Ashwin Machanavajjhala, Cheryl J. Flynn, Divesh Srivastava
CCS3
2016 Deconstructing Domain Names to Reveal Latent Topics
abstract
Measurement of the lexical properties of domain names enables many types of relatively fast, lightweight web mining analyses. These include unsupervised learning tasks such as automatic categorization and clustering of websites, as well as supervised learning tasks, such as classifying websites as malicious or benign. In this paper we explore whether these tasks can be better accomplished by identifying semantically coherent groups of words in a large set of domain names using a combination of word segmentation and topic modeling methods. By segmenting domain names to generate a large set of new domain-level features, we compare three different unsupervised learning methods for identifying topics among domain name keywords: spherical k-means clustering (SKM), Latent Dirichlet Allocation (LDA), and the Biterm Topic Model (BTM). We successfully infer semantically coherent groups of words in two independent data sets, finding that BTM topics are quantitatively the most coherent. Using the BTM, we compare inferred topics across data sets and across time periods, and we also highlight instances of homophony within the topics. Finally, we show that the BTM topics can be used as features to improve the interpretability of a supervised learning model for the detection of malicious domain names. To our knowledge this is the first large-scale empirical analysis of the co-occurrence patterns of words within domain names.
Cheryl J. Flynn, Kenneth E. Shirley
DSAA1