VLDB 2026 Research / reviewers in the wild / expert
Jordi Soria-Comas
dblp:118/0738
· DBLP profile ↗
28ranked-venue papers
15as first author
2since 2021 · last 2022
0000-0003-4112-8417ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-authorSecurity and privacy · 8 · 3 first-authorDatabases, data management, data science and information retrieval · 8 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
8 papers |
Privacy and data protection · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 60% Data stream processing · 40% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection
anonymization |
2.0 | 6 | 2022 | Multi-Dimensional Randomized Response · ICDE 2022 µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020 Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019 |
Privacy and data protection › differential privacy › local differential privacy
randomized response |
1.1 | 2 | 2022 | Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022 Multi-Dimensional Randomized Response · ICDE 2022 |
Privacy and data protection
differential privacy |
1.1 | 3 | 2022 | Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022 Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees · IEEE Trans. Inf. Forensics Secur. 2017 Enhancing data utility in differential privacy via microaggregation-based k-anonymity · VLDB J. 2014 |
Privacy and data protection › anonymization
k-anonymity |
1.0 | 4 | 2019 | Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019 t-closeness through microaggregation: Strict privacy with enhanced utility preservation · ICDE 2016 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015 |
Privacy and data protection › anonymization › microdata anonymization
microaggregation |
1.0 | 3 | 2020 | µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020 Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015 |
Privacy and data protection › differential privacy
local differential privacy |
0.6 | 1 | 2022 | Multi-Dimensional Randomized Response · ICDE 2022 |
Privacy and data protection › anonymization
t-closeness |
0.5 | 2 | 2016 | t-closeness through microaggregation: Strict privacy with enhanced utility preservation · ICDE 2016 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation · IEEE Trans. Knowl. Data Eng. 2015 |
Privacy and data protection › privacy analysis
privacy models |
0.4 | 1 | 2019 | Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams · IEEE Trans. Inf. Forensics Secur. 2019 |
Privacy and data protection › privacy-preserving data analysis
utility-preserving privacy |
0.3 | 1 | 2017 | Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees · IEEE Trans. Inf. Forensics Secur. 2017 |
Data mining
exploratory data analysis |
0.2 | 1 | 2022 | Multi-Dimensional Randomized Response · IEEE Trans. Knowl. Data Eng. 2022 |
Medical and health informatics
electronic health records |
0.1 | 1 | 2020 | µ-ANT: semantic microaggregation-based anonymization tool · Bioinform. 2020 |
Methods — techniques the papers use, named apart from their topics
clustering of attributes · 1.7microaggregation · 1.4randomized response · 1.1synthetic data generation · 0.8distribution estimation · 0.6noise-addition mechanism · 0.3suppression · 0.2generalization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Multi-Dimensional Randomized ResponseabstractIn our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods. Josep Domingo-Ferrer, Jordi Soria-Comas |
ICDE | 2 |
| 2022 | Multi-Dimensional Randomized ResponseabstractIn our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods. Josep Domingo-Ferrer, Jordi Soria-Comas |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | µ-ANT: semantic microaggregation-based anonymization toolabstractMOTIVATION: Detailed patient data are crucial for medical research. Yet, these healthcare data can only be released for secondary use if they have undergone anonymization. RESULTS: We present and describe µ-ANT, a practical and easily configurable anonymization tool for (healthcare) data. It implements several state-of-the-art methods to offer robust privacy guarantees and preserve the utility of the anonymized data as much as possible. µ-ANT also supports the heterogenous attribute types commonly found in electronic healthcare records and targets both practitioners and software developers interested in data anonymization. AVAILABILITY AND IMPLEMENTATION: (source code, documentation, executable, sample datasets and use case examples) https://github.com/CrisesUrv/microaggregation-based_anonymization_tool. David Sánchez 0001, Sergio Martínez, Josep Domingo-Ferrer, Jordi Soria-Comas, Montserrat Batet |
Bioinform. | 4 |
| 2019 | Mitigating the Curse of Dimensionality in Data Anonymization
Jordi Soria-Comas, Josep Domingo-Ferrer |
MDAI | 1 |
| 2019 | Efficient Near-Optimal Variable-Size Microaggregation
Jordi Soria-Comas, Josep Domingo-Ferrer, Rafael Mulero-Vellido |
MDAI | 1 |
| 2019 | Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data StreamsabstractAs data grow in quantity and complexity, data anonymization is becoming increasingly challenging. On one side, a great diversity of masking methods, synthetic data generation methods, and privacy models exists, and this diversity is often perceived as unsettling by practitioners. On the other side, most of the anonymization methodology was designed for static, structured, and small data, whereas the current landscape includes big data and, in particular, data streams. We explore here a unified and conceptually simple anonymization approach, by presenting a primitive called steered microaggregation that can be tailored to enforce various privacy models on static data sets and also on data streams. Steered microaggregation is based on adding artificial attributes that are properly initialized and weighted in order to guide the microaggregation process into meeting certain desired constraints. To demonstrate the potential of this type of microaggregation, we show how it can be used to achieve kanonymity, t-closeness, l-diversity, and E-differential privacy in the context of static data sets; furthermore, we discuss how it can be used to achieve k-anonymity of data streams while controlling tuple reordering. Beyond its flexibility and theoretical appeal, steered microaggregation can drastically reduce information loss, as shown by our experimental evaluation. Josep Domingo-Ferrer, Jordi Soria-Comas, Rafael Mulero-Vellido |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Anonymization of Unstructured Data via Named-Entity Recognition
Fadi Hassan, Josep Domingo-Ferrer, Jordi Soria-Comas |
MDAI | 3 |
| 2018 | Multiparty Computation with Statistical Input Confidentiality via Randomized Response
Josep Domingo-Ferrer, Rafael Mulero-Vellido, Jordi Soria-Comas |
PSD | 3 |
| 2018 | Differentially private data publishing via optimal univariate microaggregation and record perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer |
Knowl. Based Syst. | 1 |
| 2017 | A Non-Parametric Model for Accurate and Provably Private Synthetic Data SetsabstractGenerating synthetic data is a well-known option to limit disclosure risk in sensitive data releases. The usual approach is to build a model for the population and then generate a synthetic data set solely based on the model. We argue that building an accurate population model is difficult and we propose instead to approximate the original data as closely as privacy constraints permit. To enforce an ex ante privacy level when generating synthetic data, we introduce a new privacy model called ϵ synthetic privacy. Then, we describe a synthetic data generation method that satisfies ϵ-synthetic privacy. Finally, we evaluate the utility of the synthetic data generated with our method. Jordi Soria-Comas, Josep Domingo-Ferrer |
ARES | 1 |
| 2017 | A Methodology to Compare Anonymization Methods Regarding Their Risk-Utility Trade-off
Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas |
MDAI | 3 |
| 2017 | Differentially Private Data Sets Based on Microaggregation and Record Perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer |
MDAI | 1 |
| 2017 | Co-Utility: Self-Enforcing protocols for the mutual benefit of participants
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas |
Eng. Appl. Artif. Intell. | 4 |
| 2017 | Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy GuaranteesabstractDifferential privacy is a popular privacy model within the research community because of the strong privacy guarantee it offers, namely that the presence or absence of any individual in a data set does not significantly influence the results of analyses on the data set. However, enforcing this strict guarantee in practice significantly distorts data and/or limits data uses, thus diminishing the analytical utility of the differentially private results. In an attempt to address this shortcoming, several relaxations of differential privacy have been proposed that trade off privacy guarantees for improved data utility. In this paper, we argue that the standard formalization of differential privacy is stricter than required by the intuitive privacy guarantee it seeks. In particular, the standard formalization requires indistinguishability of results between any pair of neighbor data sets, while indistinguishability between the actual data set and its neighbor data sets should be enough. This limits the data controller's ability to adjust the level of protection to the actual data, hence resulting in significant accuracy loss. In this respect, we propose individual differential privacy, an alternative differential privacy notion that offers the same privacy guarantees as standard differential privacy to individuals (even though not to groups of individuals). This new notion allows the data controller to adjust the distortion to the actual data set, which results in less distortion and more analytical accuracy. We propose several mechanisms to attain individual differential privacy and we compare the new notion against standard differential privacy in terms of the accuracy of the analytical results. Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, David Megías 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | t-closeness through microaggregation: Strict privacy with enhanced utility preservationabstractThis paper proposes and shows how to use microaggregation to attain t-closeness on top of k-anonymity to protect data releases. The advantages in terms of data utility preservation of microaggregation over classic approaches based on generalizing values are analyzed. Then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated. Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez |
ICDE | 1 |
| 2016 | Anonymization in the Time of Big Data
Josep Domingo-Ferrer, Jordi Soria-Comas |
PSD | 2 |
| 2016 | Big Data Privacy: Challenges to Privacy Principles and ModelsabstractAbstract This paper explores the challenges raised by big data in privacy-preserving data management. First, we examine the conflicts raised by big data with respect to preexisting concepts of private data management, such as consent, purpose limitation, transparency and individual rights of access, rectification and erasure. Anonymization appears as the best tool to mitigate such conflicts, and it is best implemented by adhering to a privacy model with precise privacy guarantees. For this reason, we evaluate how well the two main privacy models used in anonymization (k-anonymity and $$\varepsilon $$ ε -differential privacy) meet the requirements of big data, namely composability, low computational cost and linkability. Jordi Soria-Comas, Josep Domingo-Ferrer |
Data Sci. Eng. | 1 |
| 2016 | Self-enforcing protocols via co-utile reputation management
Josep Domingo-Ferrer, Oriol Farràs, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas |
Inf. Sci. | 5 |
| 2015 | Co-utile Collaborative Anonymization of Microdata
Jordi Soria-Comas, Josep Domingo-Ferrer |
MDAI | 1 |
| 2015 | Disclosure risk assessment via record linkage by a maximum-knowledge attackerabstractBefore releasing an anonymized data set, the data protector must know how safe the data set is, that is, how much disclosure risk is incurred by the release. If no privacy model is used to select specific privacy guarantees prior to anonymization, posterior disclosure risk assessment must be performed based on the anonymized data set and, if the result is not satisfactory, anonymization must be repeated with stricter privacy parameters. Even if a privacy model is used, it may still be advisable to empirically evaluate disclosure on the anonymized data set, especially if the privacy model parameters have been relaxed to improve data utility. Record linkage is a general methodology to posterior disclosure risk assessment, whereby the data protector attempts to recreate the attacker's re-identification scenario. An important limitation of record linkage is that it usually requires the data protector to make restrictive assumptions on the attacker's background knowledge. To overcome this limitation, we present a maximum-knowledge attacker model and then we specify and compare several record linkage tests for such a worst-case attacker. Our tests are based on comparing the distribution of linkage distances between the original and the anonymized data set with the distribution of distances between one of the two previous data sets and one random data set. The more similar the distributions, the more plausibly deniable are record linkages claimed by an attacker. Because attaining zero disclosure risk for all records is too costly in terms of utility, a less demanding alternative is presented whose goal is to reduce the maximum per-record disclosure risk. Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas |
PST | 3 |
| 2015 | From t-closeness to differential privacy and vice versa in data anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas |
Knowl. Based Syst. | 2 |
| 2015 | t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility PreservationabstractMicroaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate k-anonymous data sets, where the identity of each subject is hidden within a group of k subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers, and avoiding discretization of numerical data. k-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of k subjects is too small. To address this issue, several refinements of k-anonymity have been proposed, among which t-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate t-close data sets are based on generalization and suppression (they are extensions of k-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate k-anonymous t-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated. Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | Data Anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas |
CRiSIS | 2 |
| 2014 | Enhancing data utility in differential privacy via microaggregation-based k-anonymity
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez |
VLDB J. | 1 |
| 2013 | Differential privacy via t-closeness in data publishingabstractk-Anonymity and e-differential privacy are two main privacy models proposed within the computer science community. Whereas the former was proposed for privacy-preserving data publishing, i.e. data set anonymization, the latter initially arose in the context of interactive databases and was later extended to data publishing. We show here that t-closeness, one of the extensions of k-anonymity, can actually yieldε-differential privacy in data publishing when t =exp(ε). We detail a construction based on bucketization that realizes the previous implication; hence, as an ancillary result, we provide a new computational procedure to achieve t-closeness and ε-differential privacy in data publishing. Jordi Soria-Comas, Josep Domingo-Ferrer |
PST | 1 |
| 2013 | Optimal data-independent noise for differential privacy
Jordi Soria-Comas, Josep Domingo-Ferrer |
Inf. Sci. | 1 |
| 2012 | Probabilistic k-anonymity through microaggregation and data swappingabstractk-Anonymity is a privacy property used to limit the risk of re-identification in a microdata set. A data set satisfying k-anonymity consists of groups of k records which are indistinguishable as far as their quasi-identifier attributes are concerned. Hence, the probability of re-identifying a record within a group is 1/k. We introduce the probabilistic k-anonymity property, which relaxes the indistinguishability requirement of k-anonymity and only requires that the probability of re-identification be the same as in k-anonymity. Two computational heuristics to achieve probabilistic k-anonymity based on data swapping are proposed: MDAV microaggregation on the quasi-identifiers plus swapping, and individual ranking microaggregation on individual confidential attributes plus swapping. We report experimental results, where we compare the utility of original, k-anonymous and probabilistically k-anonymous data. Jordi Soria-Comas, Josep Domingo-Ferrer |
FUZZ-IEEE | 1 |
| 2012 | Sensitivity-Independent differential Privacy via Prior Knowledge RefinementabstractWe propose a new mechanism to implement differential privacy. Unlike the usual mechanism based on adding a noise whose magnitude is proportional to the sensitivity of the query function, our proposal is based on the refinement of the user's prior knowledge about the response. Our mechanism is shown to have several advantages over noise addition: it does not require complex computations, and thus it can be easily automated; it lets the user exploit her prior knowledge about the response to achieve better data quality; and it is independent of the sensitivity of the query function (although this can be a disadvantage if the sensitivity is small). Furthermore, we give a general algorithm for knowledge refinement and we show some compounding properties of our mechanism for the case of multiple queries; also, we build an interactive mechanism on top of knowledge refinement and we show that it is safe against adaptive attacks. Finally, we give a quality assessment for the responses to individual queries. Jordi Soria-Comas, Josep Domingo-Ferrer |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 1 |