Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jordi Soria-Comas

dblp:118/0738 · DBLP profile ↗
← Back
28ranked-venue papers
15as first author
2since 2021 · last 2022
0000-0003-4112-8417ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-authorSecurity and privacy · 8 · 3 first-authorDatabases, data management, data science and information retrieval · 8 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
8 papers
Privacy and data protection · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 60% Data stream processing · 40%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Privacy and data protection
anonymization
2.062022
Multi-Dimensional Randomized Response · ICDE 2022
µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020
Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019
Privacy and data protection › differential privacy › local differential privacy
randomized response
1.122022
Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022
Multi-Dimensional Randomized Response · ICDE 2022
Privacy and data protection
differential privacy
1.132022
Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022
Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees · IEEE Trans. Inf. Forensics Secur. 2017
Enhancing data utility in differential privacy via microaggregation-based k-anonymity · VLDB J. 2014
Privacy and data protection › anonymization
k-anonymity
1.042019
Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019
t-closeness through microaggregation: Strict privacy with enhanced utility preservation · ICDE 2016
t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015
Privacy and data protection › anonymization › microdata anonymization
microaggregation
1.032020
µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020
Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019
t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015
Privacy and data protection › differential privacy
local differential privacy
0.612022
Multi-Dimensional Randomized Response · ICDE 2022
Privacy and data protection › anonymization
t-closeness
0.522016
t-closeness through microaggregation: Strict privacy with enhanced utility preservation · ICDE 2016
t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015
Privacy and data protection › privacy analysis
privacy models
0.412019
Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019
Privacy and data protection › privacy-preserving data analysis
utility-preserving privacy
0.312017
Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees · IEEE Trans. Inf. Forensics Secur. 2017
Data mining
exploratory data analysis
0.212022
Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022
Medical and health informatics
electronic health records
0.112020
µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020

Methods — techniques the papers use, named apart from their topics

clustering of attributes · 1.7microaggregation · 1.4randomized response · 1.1synthetic data generation · 0.8distribution estimation · 0.6noise-addition mechanism · 0.3suppression · 0.2generalization · 0.2
YearPublicationVenuePosition
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
ICDE2
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
IEEE Trans. Knowl. Data Eng.2
2020 µ-ANT: semantic microaggregation-based anonymization tool
abstract
MOTIVATION: Detailed patient data are crucial for medical research. Yet, these healthcare data can only be released for secondary use if they have undergone anonymization. RESULTS: We present and describe µ-ANT, a practical and easily configurable anonymization tool for (healthcare) data. It implements several state-of-the-art methods to offer robust privacy guarantees and preserve the utility of the anonymized data as much as possible. µ-ANT also supports the heterogenous attribute types commonly found in electronic healthcare records and targets both practitioners and software developers interested in data anonymization. AVAILABILITY AND IMPLEMENTATION: (source code, documentation, executable, sample datasets and use case examples) https://github.com/CrisesUrv/microaggregation-based_anonymization_tool.
David Sánchez 0001, Sergio Martínez, Josep Domingo-Ferrer, Jordi Soria-Comas, Montserrat Batet
Bioinform.4
2019 Mitigating the Curse of Dimensionality in Data Anonymization
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI1
2019 Efficient Near-Optimal Variable-Size Microaggregation
Jordi Soria-Comas, Josep Domingo-Ferrer, Rafael Mulero-Vellido
MDAI1
2019 Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams
abstract
As data grow in quantity and complexity, data anonymization is becoming increasingly challenging. On one side, a great diversity of masking methods, synthetic data generation methods, and privacy models exists, and this diversity is often perceived as unsettling by practitioners. On the other side, most of the anonymization methodology was designed for static, structured, and small data, whereas the current landscape includes big data and, in particular, data streams. We explore here a unified and conceptually simple anonymization approach, by presenting a primitive called steered microaggregation that can be tailored to enforce various privacy models on static data sets and also on data streams. Steered microaggregation is based on adding artificial attributes that are properly initialized and weighted in order to guide the microaggregation process into meeting certain desired constraints. To demonstrate the potential of this type of microaggregation, we show how it can be used to achieve kanonymity, t-closeness, l-diversity, and E-differential privacy in the context of static data sets; furthermore, we discuss how it can be used to achieve k-anonymity of data streams while controlling tuple reordering. Beyond its flexibility and theoretical appeal, steered microaggregation can drastically reduce information loss, as shown by our experimental evaluation.
Josep Domingo-Ferrer, Jordi Soria-Comas, Rafael Mulero-Vellido
IEEE Trans. Inf. Forensics Secur.2
2018 Anonymization of Unstructured Data via Named-Entity Recognition
Fadi Hassan, Josep Domingo-Ferrer, Jordi Soria-Comas
MDAI3
2018 Multiparty Computation with Statistical Input Confidentiality via Randomized Response
Josep Domingo-Ferrer, Rafael Mulero-Vellido, Jordi Soria-Comas
PSD3
2018 Differentially private data publishing via optimal univariate microaggregation and record perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer
Knowl. Based Syst.1
2017 A Non-Parametric Model for Accurate and Provably Private Synthetic Data Sets
abstract
Generating synthetic data is a well-known option to limit disclosure risk in sensitive data releases. The usual approach is to build a model for the population and then generate a synthetic data set solely based on the model. We argue that building an accurate population model is difficult and we propose instead to approximate the original data as closely as privacy constraints permit. To enforce an ex ante privacy level when generating synthetic data, we introduce a new privacy model called ϵ synthetic privacy. Then, we describe a synthetic data generation method that satisfies ϵ-synthetic privacy. Finally, we evaluate the utility of the synthetic data generated with our method.
Jordi Soria-Comas, Josep Domingo-Ferrer
ARES1
2017 A Methodology to Compare Anonymization Methods Regarding Their Risk-Utility Trade-off
Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas
MDAI3
2017 Differentially Private Data Sets Based on Microaggregation and Record Perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI1
2017 Co-Utility: Self-Enforcing protocols for the mutual benefit of participants
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Eng. Appl. Artif. Intell.4
2017 Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees
abstract
Differential privacy is a popular privacy model within the research community because of the strong privacy guarantee it offers, namely that the presence or absence of any individual in a data set does not significantly influence the results of analyses on the data set. However, enforcing this strict guarantee in practice significantly distorts data and/or limits data uses, thus diminishing the analytical utility of the differentially private results. In an attempt to address this shortcoming, several relaxations of differential privacy have been proposed that trade off privacy guarantees for improved data utility. In this paper, we argue that the standard formalization of differential privacy is stricter than required by the intuitive privacy guarantee it seeks. In particular, the standard formalization requires indistinguishability of results between any pair of neighbor data sets, while indistinguishability between the actual data set and its neighbor data sets should be enough. This limits the data controller's ability to adjust the level of protection to the actual data, hence resulting in significant accuracy loss. In this respect, we propose individual differential privacy, an alternative differential privacy notion that offers the same privacy guarantees as standard differential privacy to individuals (even though not to groups of individuals). This new notion allows the data controller to adjust the distortion to the actual data set, which results in less distortion and more analytical accuracy. We propose several mechanisms to attain individual differential privacy and we compare the new notion against standard differential privacy in terms of the accuracy of the analytical results.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, David Megías 0001
IEEE Trans. Inf. Forensics Secur.1
2016 t-closeness through microaggregation: Strict privacy with enhanced utility preservation
abstract
This paper proposes and shows how to use microaggregation to attain t-closeness on top of k-anonymity to protect data releases. The advantages in terms of data utility preservation of microaggregation over classic approaches based on generalizing values are analyzed. Then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
ICDE1
2016 Anonymization in the Time of Big Data
Josep Domingo-Ferrer, Jordi Soria-Comas
PSD2
2016 Big Data Privacy: Challenges to Privacy Principles and Models
abstract
Abstract This paper explores the challenges raised by big data in privacy-preserving data management. First, we examine the conflicts raised by big data with respect to preexisting concepts of private data management, such as consent, purpose limitation, transparency and individual rights of access, rectification and erasure. Anonymization appears as the best tool to mitigate such conflicts, and it is best implemented by adhering to a privacy model with precise privacy guarantees. For this reason, we evaluate how well the two main privacy models used in anonymization (k-anonymity and $$\varepsilon $$ ε -differential privacy) meet the requirements of big data, namely composability, low computational cost and linkability.
Jordi Soria-Comas, Josep Domingo-Ferrer
Data Sci. Eng.1
2016 Self-enforcing protocols via co-utile reputation management
Josep Domingo-Ferrer, Oriol Farràs, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Inf. Sci.5
2015 Co-utile Collaborative Anonymization of Microdata
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI1
2015 Disclosure risk assessment via record linkage by a maximum-knowledge attacker
abstract
Before releasing an anonymized data set, the data protector must know how safe the data set is, that is, how much disclosure risk is incurred by the release. If no privacy model is used to select specific privacy guarantees prior to anonymization, posterior disclosure risk assessment must be performed based on the anonymized data set and, if the result is not satisfactory, anonymization must be repeated with stricter privacy parameters. Even if a privacy model is used, it may still be advisable to empirically evaluate disclosure on the anonymized data set, especially if the privacy model parameters have been relaxed to improve data utility. Record linkage is a general methodology to posterior disclosure risk assessment, whereby the data protector attempts to recreate the attacker's re-identification scenario. An important limitation of record linkage is that it usually requires the data protector to make restrictive assumptions on the attacker's background knowledge. To overcome this limitation, we present a maximum-knowledge attacker model and then we specify and compare several record linkage tests for such a worst-case attacker. Our tests are based on comparing the distribution of linkage distances between the original and the anonymized data set with the distribution of distances between one of the two previous data sets and one random data set. The more similar the distributions, the more plausibly deniable are record linkages claimed by an attacker. Because attaining zero disclosure risk for all records is too costly in terms of utility, a less demanding alternative is presented whose goal is to reduce the maximum per-record disclosure risk.
Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas
PST3
2015 From t-closeness to differential privacy and vice versa in data anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas
Knowl. Based Syst.2
2015 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation
abstract
Microaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate k-anonymous data sets, where the identity of each subject is hidden within a group of k subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers, and avoiding discretization of numerical data. k-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of k subjects is too small. To address this issue, several refinements of k-anonymity have been proposed, among which t-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate t-close data sets are based on generalization and suppression (they are extensions of k-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate k-anonymous t-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
IEEE Trans. Knowl. Data Eng.1
2014 Data Anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas
CRiSIS2
2014 Enhancing data utility in differential privacy via microaggregation-based k-anonymity
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
VLDB J.1
2013 Differential privacy via t-closeness in data publishing
abstract
k-Anonymity and e-differential privacy are two main privacy models proposed within the computer science community. Whereas the former was proposed for privacy-preserving data publishing, i.e. data set anonymization, the latter initially arose in the context of interactive databases and was later extended to data publishing. We show here that t-closeness, one of the extensions of k-anonymity, can actually yieldε-differential privacy in data publishing when t =exp(ε). We detail a construction based on bucketization that realizes the previous implication; hence, as an ancillary result, we provide a new computational procedure to achieve t-closeness and ε-differential privacy in data publishing.
Jordi Soria-Comas, Josep Domingo-Ferrer
PST1
2013 Optimal data-independent noise for differential privacy
Jordi Soria-Comas, Josep Domingo-Ferrer
Inf. Sci.1
2012 Probabilistic k-anonymity through microaggregation and data swapping
abstract
k-Anonymity is a privacy property used to limit the risk of re-identification in a microdata set. A data set satisfying k-anonymity consists of groups of k records which are indistinguishable as far as their quasi-identifier attributes are concerned. Hence, the probability of re-identifying a record within a group is 1/k. We introduce the probabilistic k-anonymity property, which relaxes the indistinguishability requirement of k-anonymity and only requires that the probability of re-identification be the same as in k-anonymity. Two computational heuristics to achieve probabilistic k-anonymity based on data swapping are proposed: MDAV microaggregation on the quasi-identifiers plus swapping, and individual ranking microaggregation on individual confidential attributes plus swapping. We report experimental results, where we compare the utility of original, k-anonymous and probabilistically k-anonymous data.
Jordi Soria-Comas, Josep Domingo-Ferrer
FUZZ-IEEE1
2012 Sensitivity-Independent differential Privacy via Prior Knowledge Refinement
abstract
We propose a new mechanism to implement differential privacy. Unlike the usual mechanism based on adding a noise whose magnitude is proportional to the sensitivity of the query function, our proposal is based on the refinement of the user's prior knowledge about the response. Our mechanism is shown to have several advantages over noise addition: it does not require complex computations, and thus it can be easily automated; it lets the user exploit her prior knowledge about the response to achieve better data quality; and it is independent of the sensitivity of the query function (although this can be a disadvantage if the sensitivity is small). Furthermore, we give a general algorithm for knowledge refinement and we show some compounding properties of our mechanism for the case of multiple queries; also, we build an interactive mechanism on top of knowledge refinement and we show that it is safe against adaptive attacks. Finally, we give a quality assessment for the responses to individual queries.
Jordi Soria-Comas, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1