Josep Domingo-Ferrer

dblp:d/JDomingoFerrer · DBLP profile ↗
← Back
213ranked-venue papers
77as first author
37since 2021 · last 2026
0000-0001-7213-4962ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 68 · 26 first-author · 10 since 2021Artificial intelligence and machine learning · 59 · 17 first-author · 16 since 2021Databases, data management, data science and information retrieval · 50 · 25 first-author · 7 since 2021Computer networks · 22 · 9 first-author · 4 since 2021Systems, architecture and hardware · 6 · 2 first-authorTheory of computation · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 FedVendorBC: Accurate Multi-vendor Federated Learning for Privacy-Preserving and Generalizable Breast Cancer Diagnosis Using Ultrasound Images
David Sánchez 0001, Zouhair Haddi, Josep Domingo-Ferrer
AIME (1)4
2026 Explainability-Driven Image Anonymization in Latent Space (EDIALS)
abstract
Facial image anonymization is essential to enable privacy-preserving image data sharing. The core challenge lies in removing identity-revealing information without degrading the utility of the images, which is essential for, e.g., demographic analysis. However, existing techniques apply uniform pixel-level distortions or synthesize replacements using Generative Adversarial Networks (GANs), which do not retain the meaningful features necessary for downstream tasks. To address this issue, we introduce EDIALS (Explainability-Driven Image Anonymization in Latent Space), which selectively modifies identity-specific latent features identified via explainability techniques. By applying targeted, incremental distortions in the latent space of an adversarial autoencoder, EDIALS effectively anonymizes images while preserving their analytical utility much better than existing techniques. Empirical evaluations on a common dataset show that EDIALS achieves 0.42% re-identification risk (equivalent to random guessing) while maintaining high utility: 84.66% F1 for age, 97.91% for gender, and 82.61% for race classification. In contrast, DeepPrivacy2 —a state-of-the-art GAN-based approach— results in a re-identification risk as large as 16.27% and lower utility: 78.32% F1, 82.58%, and 73.32% for age, gender, and race classification, respectively.
Younas Khan, Anna Monreale, Carlo Metta, David Sánchez 0001, Josep Domingo-Ferrer
CODASPY5
2026 Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions
abstract
Membership inference attacks (MIAs) have become the standard tool for evaluating privacy leakage in machine learning (ML). Among them, the Likelihood-Ratio Attack (LiRA) is widely regarded as the state of the art when sufficient shadow models are available. However, prior evaluations have often overstated the effectiveness of LiRA by attacking models overconfident on their training samples, calibrating thresholds on target data, assuming balanced membership priors, and/or overlooking attack reproducibility. We re-evaluate LiRA under a realistic protocol that (i) trains models using anti-overfitting (AOF) (and transfer learning (TL), when applicable) to reduce overconfidence as it would be desirable in production models; (ii) calibrates decision thresholds from shadow models and data rather than (usually unavailable) target data; (iii) measures positive predictive value (PPV, a.k.a. precision) under shadow-based thresholds and skewed - rather than unrealistically balanced - membership priors (pi <= 10%); and (iv) quantifies per-sample membership reproducibility across different seeds and training variations. In this setting, we find that (a) AOF significantly weakens LiRA and TL further reduces the effectiveness of the attack, while improving model accuracy; (b) with shadow-based thresholds and skewed priors, LiRA's PPV often drops from near-perfect to substantially lower levels, especially under AOF/AOF+TL and for pi <= 10%; and (c) LiRA's thresholded vulnerable sets at extremely low FPR exhibit poor reproducibility across runs, while likelihood ratio-based rankings are more stable. These results suggest that (i) LiRA, and likely weaker MIAs, are less effective than previously suggested, and their positive inferences can be less reliable under realistic settings; and (ii) for MIAs to serve as meaningful privacy auditing tools, their evaluation must reflect pragmatic training practices, feasible attacker assumptions, and reproducibility considerations. We release our code at: https://github.com/najeebjebreel/lira_analysis.
Najeeb Jebreel, Mona Khalil, David Sánchez 0001, Josep Domingo-Ferrer
Proc. Priv. Enhancing Technol.4
2025 Defenses Against Membership Inference Attacks on Unlearned Data
Josep Domingo-Ferrer, Najeeb Jebreel, David Sánchez 0001
MDAI1
2025 Privacy Risks in Machine Learning: Truths and Myths
Josep Domingo-Ferrer
SECRYPT1
2025 Enhancing efficiency and data utility in longitudinal data anonymization
Fatemeh Amiri, David Sánchez 0001, Josep Domingo-Ferrer
Inf. Sci.3
2025 MemberShield: A framework for federated learning with membership privacy
abstract
Federated Learning (FL) allows multiple data owners to build high-quality deep learning models collaboratively, by sharing only model updates and keeping data on their premises. Even though FL offers privacy-by-design, it is vulnerable to membership inference attacks (MIA), where an adversary tries to determine whether a sample was included in the training data. Existing defenses against MIA cannot offer meaningful privacy protection without significantly hampering the model's utility and causing a non-negligible training overhead. In this paper we analyze the underlying causes of the differences in the model behavior for member and non-member samples, which arise from model overfitting and facilitate MIAs. Accordingly, we propose MemberShield, a generalization-based defense method for MIAs that consists of: (i) one-time preprocessing of each client's training data labels that transforms one-hot encoded labels to soft labels and eventually exploits them in local training, and (ii) early stopping the training when the local model's validation accuracy does not improve on that of the global model for a number of epochs. Extensive empirical evaluations on three widely used datasets and four model architectures demonstrate that MemberShield outperforms state-of-the-art defense methods by delivering substantially better practical privacy protection against all forms of MIAs, while better preserving the target model utility. On top of that, our proposal significantly reduces training time and is straightforward to implement, by just tuning a single hyperparameter.
David Sánchez 0001, Zouhair Haddi, Josep Domingo-Ferrer
Neural Networks4
2025 DP2Unlearning: An efficient and guaranteed unlearning framework for LLMs
abstract
Large language models (LLMs) have recently revolutionized language processing tasks but have also brought ethical and legal issues. LLMs have a tendency to memorize potentially private or copyrighted information present in the training data, which might then be delivered to end users at inference time. When this happens, a naive solution is to retrain the model from scratch after excluding the undesired data. Although this guarantees that the target data have been forgotten, it is also prohibitively expensive for LLMs. Approximate unlearning offers a more efficient alternative, as it consists of ex post modifications of the trained model itself to prevent undesirable results, but it lacks forgetting guarantees because it relies solely on empirical evidence. In this work, we present DP2Unlearning, a novel LLM unlearning framework that offers formal forgetting guarantees at a significantly lower cost than retraining from scratch on the data to be retained. DP2Unlearning involves training LLMs on textual data protected using ϵ-differential privacy (DP), which later enables efficient unlearning with the guarantees against disclosure associated with the chosen ϵ. Our experiments demonstrate that DP2Unlearning achieves similar model performance post-unlearning, compared to an LLM retraining from scratch on retained data -the gold standard exact unlearning- but at approximately half the unlearning cost. In addition, with a reasonable computational cost, it outperforms approximate unlearning methods at both preserving the utility of the model post-unlearning and effectively forgetting the targeted information. The code of our experiments is available at https://github.com/tamimalmahmud/LLM-Unlearning/tree/main/DP2Unlearning.
Tamim Al Mahmud, Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001
Neural Networks3
2025 Synthetic Data Generation via the Permutation Paradigm With Optional $k$k-Anonymity
abstract
Most methods in the literature on synthetic microdata (individual records) generation are parametric, that is, they require knowing or estimating the joint or the conditional distribution of the original microdata. This may be a significant hurdle unless the original microdata are multivariate normal. We propose a rank-based approach to generating synthetic microdata based on the permutation paradigm. We present three different methods and we analyze the utility and the confidentiality they afford. The third method is actually an extension of the second method that adds$k$-anonymity protection against reidentification to the confidentiality against attribute disclosure offered by the first two methods. Our algorithms only require the identification of themarginaldistributions of attributes and yield synthetic attributes that replicate the relationships between the original attributes exclusively based on ranks. This proposal is especially attractive for non-normal or multi-type microdata.
Josep Domingo-Ferrer, Krishnamurty Muralidhar, Sergio Martínez
IEEE Trans. Dependable Secur. Comput.1
2025 Dual-Server Privacy-Preserving Collaborative Deep Learning: A Round-Efficient, Dynamic and Lossless Approach
abstract
To address limitations in existing privacy-preserving collaborative deep learning (CDL) schemes, we propose a dual-server privacy-preserving CDL scheme based on homomorphic encryption and amasking technique. Specifically, in our scheme a random seed is used to initialize a pseudorandom generator that produces multiple pseudorandom numbers. These pseudorandom numbers, along with a random noise, are utilized to generate masks that are added to all parameters of a participant's locally trained model. By using homomorphic encryption, the random noise can be encrypted and eventually used to remove the masks with low message expansion. This also ensures that the global model is lossless in accuracy. Furthermore, if participants join or leave the system, only the time required to complete both model update aggregation and encrypted masks aggregation is affected. We demonstrate that our scheme is round-efficient, dynamic and lossless. We also show that it is secure against inference attacks and can resist collusion attacks of up to$t-2$participants and one of the two servers, where$t$is a security parameter indicating the minimum number of participants that participate in an aggregation round.
Lulu Wang 0015, Lei Zhang 0009, Kim-Kwang Raymond Choo, Josep Domingo-Ferrer, Mauro Conti
IEEE Trans. Dependable Secur. Comput.4
2024 Defending Against Backdoor Attacks by Layer-wise Feature Analysis (Extended Abstract)
Najeeb Jebreel, Josep Domingo-Ferrer, Yiming Li 0004
IJCAI2
2024 Machine Learning in Finance
abstract
This workshop aims to explore the intersection of Generative AI with the rich tapestry of financial data types, seeking to uncover new methodologies and techniques that can enhance predictive analytics, fraud detection, and customer insights across the sector. By harnessing these advancements in AI, we can pave the way to not only understand customer behavior but also anticipate their needs more effectively, leading to superior customer outcomes and more personalized services. Our objective is to shed light on the challenges and opportunities presented by the diverse data formats in finance. We aim to bridge the gap between the dominance of traditional models for tabular data analysis and the emerging potential of Generative AI to revolutionize the treatment of time series, click streams, and other unstructured data forms.
Leman Akoglu, Nitesh V. Chawla, Josep Domingo-Ferrer, Eren Kurshan, Senthil Kumar, Vidyut M. Naware, José A. Rodríguez-Serrano, Isha Chaturvedi, Saurabh Nagrecha, Mahashweta Das, Tanveer A. Faruquie
KDD3
2024 An Examination of the Alleged Privacy Threats of Confidence-Ranked Reconstruction of Census Microdata
David Sánchez 0001, Najeeb Jebreel, Krishnamurty Muralidhar, Josep Domingo-Ferrer, Alberto Blanco-Justicia
PSD4
2024 LFighter: Defending against the label-flipping attack in federated learning
Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001, Alberto Blanco-Justicia
Neural Networks2
2024 Enhanced Security and Privacy via Fragmented Federated Learning
abstract
In federated learning (FL), a set of participants share updates computed on their local data with an aggregator server that combines updates into a global model. However, reconciling accuracy with privacy and security is a challenge to FL. On the one hand, good updates sent by honest participants may reveal their private local information, whereas poisoned updates sent by malicious participants may compromise the model's availability and/or integrity. On the other hand, enhancing privacy via update distortion damages accuracy, whereas doing so via update aggregation damages security because it does not allow the server to filter out individual poisoned updates. To tackle the accuracy-privacy-security conflict, we propose fragmented FL (FFL), in which participants randomly exchange and mix fragments of their updates before sending them to the server. To achieve privacy, we design a lightweight protocol that allows participants to privately exchange and mix encrypted fragments of their updates so that the server can neither obtain individual updates nor link them to their originators. To achieve security, we design a reputation-based defense tailored for FFL that builds trust in participants and their mixed updates based on the quality of the fragments they exchange and the mixed updates they send. Since the exchanged fragments' parameters keep their original coordinates and attackers can be neutralized, the server can correctly reconstruct a global model from the received mixed updates without accuracy loss. Experiments on four real data sets show that FFL can prevent semi-honest servers from mounting privacy attacks, can effectively counter-poisoning attacks, and can keep the accuracy of the global model.
Najeeb Jebreel, Josep Domingo-Ferrer, Alberto Blanco-Justicia, David Sánchez 0001
IEEE Trans. Neural Networks Learn. Syst.2
2024 Dual-Server-Based Lightweight Privacy-Preserving Federated Learning
abstract
Federated learning (FL) allows multiple users to collaboratively train global machine learning models by keeping their data sets local. However, the existing privacy-preserving FL schemes suffer from several limitations, e.g., loss of accuracy, high communication/computation cost, failure to support dynamic users, and insecurity against collusion attacks. To solve these limitations, we propose a lightweight privacy-preserving FL scheme based on a dual-server architecture. Our scheme involves only lightweight cryptographic operations, i.e., hash and symmetric encryption operations, and it has low communication overhead. Thus, it is computationally lightweight and round-efficient. Further, it allows users to join/quit an FL task and it is accuracy-lossless. We formally prove that our scheme remains secure even in case of collusion attacks. In particular, if an attacker colludes with one of the servers and all the users who participate in an FL task except two, the privacy of user gradients stays unviolated. The reported experimental results demonstrate that our scheme incurs only a marginal increase in total communication overhead compared to the FL scheme without any privacy protection. In terms of computation overhead, the cost per user remains stable as the number of users grows, while the cost for the server is comparable to that of the FL scheme without any privacy protection.
Liangyu Zhong, Lulu Wang 0015, Lei Zhang 0009, Josep Domingo-Ferrer, Lin Xu 0010, Changti Wu
IEEE Trans. Netw. Serv. Manag.4
2023 Defending Against Backdoor Attacks by Layer-wise Feature Analysis
Najeeb Jebreel, Josep Domingo-Ferrer, Yiming Li 0004
PAKDD (2)2
2023 Explaining predictions and attacks in federated learning via random forests
abstract
Abstract Artificial intelligence (AI) is used for various purposes that are critical to human life. However, most state-of-the-art AI algorithms are black-box models, which means that humans cannot understand how such models make decisions. To forestall an algorithm-based authoritarian society, decisions based on machine learning ought to inspire trust by being explainable . For AI explainability to be practical, it must be feasible to obtain explanations systematically and automatically. A usual methodology to explain predictions made by a (black-box) deep learning model is to build a surrogate model based on a less difficult, more understandable decision algorithm. In this work, we focus on explaining by means of model surrogates the (mis)behavior of black-box models trained via federated learning. Federated learning is a decentralized machine learning technique that aggregates partial models trained by a set of peers on their own private data to obtain a global model. Due to its decentralized nature, federated learning offers some privacy protection to the participating peers. Nonetheless, it remains vulnerable to a variety of security attacks and even to sophisticated privacy attacks. To mitigate the effects of such attacks, we turn to the causes underlying misclassification by the federated model, which may indicate manipulations of the model. Our approach is to use random forests containing decision trees of restricted depth as surrogates of the federated black-box model. Then, we leverage decision trees in the forest to compute the importance of the features involved in the wrong predictions. We have applied our method to detect security and privacy attacks that malicious peers or the model manager may orchestrate in federated learning scenarios. Empirical results show that our method can detect attacks with high accuracy and, unlike other attack detection mechanisms, it can also explain the operation of such attacks at the peers’ side.
Rami Haffar, David Sánchez 0001, Josep Domingo-Ferrer
Appl. Intell.3
2023 Secure, accurate and privacy-aware fully decentralized learning via co-utility
abstract
Fully decentralized learning is a setting in which each peer in a P2P network trains a machine learning model with the help of the other peers. Each peer acts as a model manager by periodically sending her current model to other peers, who answer by returning model updates they compute on their private data. This creates a tension among privacy, accuracy and security. The privacy risk is that model updates returned by a peer can leak some of the peer’s private data. Unfortunately, distorting model updates to protect privacy works against the accuracy of the trained model. On the other hand, aggregating the updates of several peers and then sending the aggregate to the model manager may preserve privacy but it goes against security, because the model manager cannot filter out individual bad updates. Also, peers are autonomous and hence it cannot be taken for granted that they will honestly supply model updates to help the model manager train her model. To reconcile accuracy, privacy and security, we present a fully decentralized learning protocol such that: (i) it allows perfectly accurate individual updates to be returned by peers to the model manager in a privacy-preserving manner; (ii) it is co-utile by design, that is, it incentivizes rational peers to follow the protocol without deviating. The latter feature discourages rational attacks that might compromise security and it also deters free riding, thereby ensuring the sustainability of the protocol.
Jesús A. Manjón, Josep Domingo-Ferrer, David Sánchez 0001, Alberto Blanco-Justicia
Comput. Commun.2
2023 Fair detection of poisoning attacks in federated learning on non-i.i.d. data
Ashneet Khandpur Singh, Alberto Blanco-Justicia, Josep Domingo-Ferrer
Data Min. Knowl. Discov.3
2023 FL-Defender: Combating targeted attacks in federated learning
Najeeb Jebreel, Josep Domingo-Ferrer
Knowl. Based Syst.2
2023 Circuit-Free General-Purpose Multi-Party Computation via Co-Utile Unlinkable Outsourcing
abstract
Multiparty computation (MPC) consists in several parties engaging in joint computation in such a way that each party's input and output remain private to that party. Whereas MPC protocols for specific computations have existed since the 1980s, only recently general-purpose compilers have been developed to allow MPC on arbitrary functions. Yet, using today's MPC compilers requires substantial programming effort and skill on the user's side, among other things because nearly all compilers translate the code of the computation into a Boolean or arithmetic circuit. In particular, the circuit representation requires unrolling loops and recursive calls, which forces programmers to (often manually) define loop bounds and hardly use recursion. We present an approach allowing MPC on an arbitrary computation expressed as ordinary code with all functionalities that does not need to be translated into a circuit. Our notion of input and output privacy is predicated on unlinkability. Our method leverages co-utile computation outsourcing using anonymous channels via decentralized reputation, makes a minimalistic use of cryptography and does not require participants to be honest-but-curious: it works as long as participants are rational (self-interested), which may include rationally malicious peers (who become attackers if this is advantageous to them). We present example applications, including e-voting. Our empirical work shows that reputation captures well the behavior of peers and ensures that parties with high reputation obtain correct results.
Josep Domingo-Ferrer, Jesús A. Manjón
IEEE Trans. Dependable Secur. Comput.1
2023 Utility-Preserving Privacy Protection of Textual Documents via Word Embeddings
abstract
A great variety of mechanisms have been proposed to protect structured databases with numerical and categorical attributes; however, little attention has been devoted to unstructured textual data. Textual data protection requires first detecting sensitive pieces of text and then masking those pieces via suppression or generalization. Current solutions rely on classifiers that can recognize a fixed set of (allegedly sensitive) named entities. Yet, such approaches fall short of providing adequate protection because in reality references to sensitive information are not limited to a predefined set of entity types, and not all the appearances of certain entity type result in disclosure. In this work we propose a more general and flexible based on the notion of word embedding. By means of word embeddings we build vectors that numerically capture the semantic relationships of the textual terms. Then we evaluate the disclosure caused by the terms on the entity to be protected according to the similarity between their vector representations. Our method also preserves the semantics (and, therefore, the utility) of the document by replacing risky terms with privacy-preserving generalizations. Empirical results show that our approach offers much more robust protection and greater utility preservation than methods based on named entity recognition.
Fadi Hassan, David Sánchez 0001, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.3
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
ICDE1
2022 Measuring Fairness in Machine Learning Models via Counterfactual Examples
Rami Haffar, Ashneet Khandpur Singh, Josep Domingo-Ferrer, Najeeb Jebreel
MDAI3
2022 Bistochastic Privacy
Nicolás Ruiz, Josep Domingo-Ferrer
MDAI2
2022 Generation of Synthetic Trajectory Microdata from Language Models
Alberto Blanco-Justicia, Najeeb Jebreel, Jesús A. Manjón, Josep Domingo-Ferrer
PSD4
2022 Tit-for-Tat Disclosure of a Binding Sequence of User Analyses in Safe Data Access Centers
Josep Domingo-Ferrer
PSD1
2022 Decentralized k-anonymization of trajectories via privacy-preserving tit-for-tat
abstract
Mobility data, and specifically trajectories, are used to monitor the mobility of the population and are crucial to improve public health, transportation, urban planning, economic planning, etc. However, trajectories are personally identifiable information and hence they should be anonymized before releasing them for secondary use. Anonymization cannot be limited to suppressing the metadata containing the subject’s identity, because the origin, the destination and even the intermediate points of a trajectory may allow re-identifying the subject who followed it. Proper anonymization requires masking detailed spatiotemporal information. The standard approach to build anonymized data sets is centralized: the subjects send their original movement data to a controller, who takes care of producing an anonymized mobility data set. This requires subjects to blindly trust the controller. In this paper, we empower subjects with the ability to anonymize their trajectories locally by adhering to a privacy model in order to achieve formal privacy guarantees. After reviewing the state of the art, we motivate our choice of k-anonymity as a privacy model. We then set out to decentralize k-anonymity in a rational setting: a subject k-anonymizes her completed trajectory by aggregating with k−1 similar trajectories obtained from other (unknown) subjects. The latter trajectories are gathered via an anonymous and privacy-preserving tit-for-tat data exchange protocol, which runs on a fully decentralized peer-to-peer network. Experiments show that, without relying on a (trusted) data controller and while ensuring privacy w.r.t. other peers, our approach yields k-anonymized mobility data sets that are still reasonably useful compared to the near-optimal data sets obtained in the centralized approach.
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001
Comput. Commun.1
2022 Generating Deep Learning Model-Specific Explanations at the End User's Side
abstract
End users who cannot afford to collect and label big data to train accurate deep learning (DL) models resort to Machine Learning as a Service (MLaaS) providers, who provide paid access to accurate DL models. However, the lack of transparency in how the providers’ models make predictions causes a problem of trust. A way to increase trust (and also to align with ethical regulations) is for predictions to be accompanied by explanations locally and independently generated by the end users (rather than by explanations offered by the model providers). Explanation methods using internal components of DL models (a.k.a. model-specific explanations) are more accurate and effective than those relying solely on the inputs and outputs (a.k.a. model-agnostic explanations). However, end users lack white-box access to the internal components of the providers’ models. To tackle this issue, we propose a novel approach allowing an end user to locally generate model-specific explanations for a DL classification model accessed via a provider’s API. First, we approximate the provider’s model with a local surrogate model. We then use the surrogate model’s components to locally generate model-specific explanations that approximate the explanations obtainable with white-box access to the provider’s DL model. Specifically, we leverage the surrogate model’s gradients to generate adversarial examples that counterfactually explain why an input example is classified into a specific class. Our approach only requires the end user to have unlabeled data of size [Formula: see text] of the provider’s training data and with a similar distribution; given the small size and unlabeled nature of these data, they can be assumed to be already available to the end user or even to be supplied by the provider to build trust in his model. We demonstrate the accuracy and effectiveness of our approach through extensive experiments on two ML tasks: image classification and tabular data classification. The locally generated explanations are consistent with those obtainable with white-box access to the provider’s model, thus giving end users an independent and reliable way to determine if the provider’s model is trustworthy.
Rami Haffar, Najeeb Jebreel, David Sánchez 0001, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.4
2022 Secure and Privacy-Preserving Federated Learning via Co-Utility
abstract
The decentralized nature of federated learning, that often leverages the power of edge devices, makes it vulnerable to attacks against privacy and security. The privacy risk for a peer is that the model update she computes on her private data may, when sent to the model manager, leak information on those private data. Even more obvious are security attacks, whereby one or several malicious peers return wrong model updates in order to disrupt the learning process and lead to a wrong model being learned. In this article, we build a federated learning framework that offers privacy to the participating peers as well as security against the Byzantine and poisoning attacks. Our framework consists of several protocols that provide strong privacy to the participating peers via unlinkable anonymity and that are rationally sustainable based on the co-utility property. In other words, no rational party is interested in deviating from the proposed protocols. We leverage the notion of co-utility to build a decentralized co-utile reputation management system that provides incentives for parties to adhere to the protocols. Unlike privacy protection via differential privacy, our approach preserves the values of model updates and, hence, the accuracy of plain federated learning; unlike privacy protection via update aggregation, our approach preserves the ability to detect bad model updates while substantially reducing the computational overhead compared to methods based on homomorphic encryption.
Josep Domingo-Ferrer, Alberto Blanco-Justicia, Jesús A. Manjón, David Sánchez 0001
IEEE Internet Things J.1
2022 Differentially private publication of database streams via hybrid video coding
abstract
While most anonymization technology available today is designed for static and small data, the current picture is of massive volumes of dynamic data arriving at unprecedented velocities. From the standpoint of anonymization, the most challenging type of dynamic data is data streams. However, while the majority of proposals deal with publishing either count-based or aggregated statistics about the underlying stream, little attention has been paid to the problem of continuously publishing the stream itself with differential privacy guarantees. In this work, we propose an anonymization method that can publish multiple numerical-attribute, finite microdata streams with high protection as well as high utility, the latter aspect measured as data distortion, delay and record reordering. Our method, which relies on the well-known differential pulse-code modulation scheme, adapts techniques originally intended for hybrid video encoding, to favor and leverage dependencies among the blocks of the original stream and thereby reduce data distortion. The proposed solution is assessed experimentally on two of the largest data sets in the scientific community working in data anonymization. Our extensive empirical evaluation shows the trade-off among privacy protection, data distortion, delay and record reordering, and demonstrates the suitability of adapting video-compression techniques to anonymize database streams.
Javier Parra-Arnau, Thorsten Strufe, Josep Domingo-Ferrer
Knowl. Based Syst.3
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
IEEE Trans. Knowl. Data Eng.1
2021 Towards Machine Learning-Assisted Output Checking for Statistical Disclosure Control
Josep Domingo-Ferrer, Alberto Blanco-Justicia
MDAI1
2021 Explaining Image Misclassification in Deep Learning via Adversarial Examples
Rami Haffar, Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001
MDAI3
2021 Achieving security and privacy in federated learning systems: Survey, research challenges and future directions
abstract
Federated learning (FL) allows a server to learn a machine learning (ML) model across multiple decentralized clients that privately store their own training data. In contrast with centralized ML approaches, FL saves computation to the server and does not require the clients to outsource their private data to the server. However, FL is not free of issues. On the one hand, the model updates sent by the clients at each training epoch might leak information on the clients’ private data. On the other hand, the model learnt by the server may be subjected to attacks by malicious clients; these security attacks might poison the model or prevent it from converging. In this paper, we first examine security and privacy attacks to FL and critically survey solutions proposed in the literature to mitigate each attack. Afterwards, we discuss the difficulty of simultaneously achieving security and privacy protection. Finally, we sketch ways to tackle this open problem and attain both security and privacy.
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Adrian Flanagan, Kuan Eeik Tan
Eng. Appl. Artif. Intell.2
2021 General Confidentiality and Utility Metrics for Privacy-Preserving Data Publishing Based on the Permutation Model
abstract
Anonymization for privacy-preserving data publishing, also known as statistical disclosure control (SDC), can be viewed under the lens of the permutation model. According to this model, any SDC method for individual data records is functionally equivalent to a permutation step plus a noise addition step, where the noise added is marginal, in the sense that it does not alter ranks. Here, we propose metrics to quantify the data confidentiality and utility achieved by SDC methods based on the permutation model. We distinguish two privacy notions: in our work, anonymity refers to subjects and hence mainly to protection against record re-identification, whereas confidentiality refers to the protection afforded to attribute values against attribute disclosure. Thus, our confidentiality metrics are useful even if using a privacy model ensuring an anonymity level ex ante. The utility metric is a general-purpose metric that can be conveniently traded off against the confidentiality metrics, because all of them are bounded between 0 and 1. As an application, we compare the utility-confidentiality trade-offs achieved by several anonymization approaches, including privacy models (k-anonymity and ε-differential privacy) as well as SDC methods (additive noise, multiplicative noise and synthetic data) used without privacy models.
Josep Domingo-Ferrer, Krishnamurty Muralidhar, Maria Bras-Amorós
IEEE Trans. Dependable Secur. Comput.1
2020 Co-Utile Peer-to-Peer Decentralized Computing
abstract
Outsourcing computation allows wielding huge computational power. Even though cloud computing is the most usual type of outsourcing, resorting to idle edge devices for decentralized computation is an increasingly attractive alternative. We tackle the problem of making peer honesty and thus computation correctness self-enforcing in decentralized computing with untrusted peers. To do so, we leverage the co-utility property, which characterizes a situation in which honest co-operation is the best rational option to take even for purely selfish agents; in particular, if a protocol is co-utile, it is self-enforcing. Reputation is a powerful incentive that can make a P2P protocol co-utile. We present a co-utile P2P decentralized computing protocol that builds on a decentralized reputation calculation, which is itself co-utile and therefore self-enforcing. In this protocol, peers are given a computational task including code and data and they are incentivized to compute it correctly. Based also on co-utile reputation, we then present a protocol for federated learning, whereby peers compute on their local private data and have no incentive to randomly attack or poison the model. Our experiments show the viability of our co-utile approach to obtain correct results in both decentralized computation and federated learning.
Josep Domingo-Ferrer, Alberto Blanco-Justicia, David Sánchez 0001, Najeeb Jebreel
CCGRID1
2020 Fair Detection of Poisoning Attacks in Federated Learning
abstract
Federated learning is a decentralized machine learning technique that aggregates partial models trained by a set of clients on their own private data to obtain a global model. This technique is vulnerable to security attacks, such as model poisoning, whereby malicious clients submit bad updates in order to prevent the model from converging or to introduce artificial bias in the classification. Applying anti-poisoning techniques might lead to the discrimination of minority groups whose data are significantly and legitimately different from those of the majority of clients. In this work, we strive to strike a balance between fighting poisoning and accommodating diversity to help learning fairer and less discriminatory federated learning models. In this way, we forestall the exclusion of diverse clients while still ensuring detection of poisoning attacks. Empirical work on a standard machine learning data set shows that employing our approach to tell legitimate from malicious updates produces models that are more accurate than those obtained with standard poisoning detection techniques.
Ashneet Khandpur Singh, Alberto Blanco-Justicia, Josep Domingo-Ferrer, David Sánchez 0001, David Rebollo-Monedero
ICTAI3
2020 Privacy-Preserving Computation of the Earth Mover's Distance
Alberto Blanco-Justicia, Josep Domingo-Ferrer
ISC2
2020 Explaining Misclassification and Attacks in Deep Learning via Random Forests
Rami Haffar, Josep Domingo-Ferrer, David Sánchez 0001
MDAI2
2020 Efficient Detection of Byzantine Attacks in Federated Learning Using Last Layer Biases
Najeeb Jebreel, Alberto Blanco-Justicia, David Sánchez 0001, Josep Domingo-Ferrer
MDAI4
2020 Detecting Bad Answers in Survey Data Through Unsupervised Machine Learning
Najeeb Jebreel, Rami Haffar, Ashneet Khandpur Singh, David Sánchez 0001, Josep Domingo-Ferrer, Alberto Blanco-Justicia
PSD5
2020 ε-Differential Privacy for Microdata Releases Does Not Guarantee Confidentiality (Let Alone Utility)
Krishnamurty Muralidhar, Josep Domingo-Ferrer, Sergio Martínez
PSD2
2020 µ-ANT: semantic microaggregation-based anonymization tool
abstract
MOTIVATION: Detailed patient data are crucial for medical research. Yet, these healthcare data can only be released for secondary use if they have undergone anonymization. RESULTS: We present and describe µ-ANT, a practical and easily configurable anonymization tool for (healthcare) data. It implements several state-of-the-art methods to offer robust privacy guarantees and preserve the utility of the anonymized data as much as possible. µ-ANT also supports the heterogenous attribute types commonly found in electronic healthcare records and targets both practitioners and software developers interested in data anonymization. AVAILABILITY AND IMPLEMENTATION: (source code, documentation, executable, sample datasets and use case examples) https://github.com/CrisesUrv/microaggregation-based_anonymization_tool.
David Sánchez 0001, Sergio Martínez, Josep Domingo-Ferrer, Jordi Soria-Comas, Montserrat Batet
Bioinform.3
2020 Outsourcing analyses on privacy-protected multivariate categorical data stored in untrusted clouds
Josep Domingo-Ferrer, David Sánchez 0001, Sara Ricci, Mónica Muñoz-Batista
Knowl. Inf. Syst.1
2020 Machine learning explainability via microaggregation and shallow decision trees
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001
Knowl. Based Syst.2
2019 18th Workshop on Privacy in the Electronic Society (WPES 2019)
abstract
The 18th Workshop on Privacy in the Electronic Society (WPES 2019) was held on November 11th, 2019, in conjunction with the 26th ACM Conference on Computer and Communications Security (CCS 2019) in London, United Kingdom. The goal of WPES is to bring together privacy researchers and practitioners to discuss the privacy problems that arise in an interconnected society and solutions to those problems. The program for the workshop contains 14 full papers and 5 short papers selected from a total of 67 submissions. Specific topics covered in the program include secure computation, secure communication, mobile-device privacy, genomic data privacy, social aspects of privacy, data anonymization, and privacy-enhancing techniques for the Internet of Things, blockchain and Web.
Josep Domingo-Ferrer
CCS1
2019 Machine Learning Explainability Through Comprehensible Decision Trees
Alberto Blanco-Justicia, Josep Domingo-Ferrer
CD-MAKE2
2019 Personal Big Data, GDPR and Anonymization
Josep Domingo-Ferrer
FQAS1
2019 Mitigating the Curse of Dimensionality in Data Anonymization
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI2
2019 Efficient Near-Optimal Variable-Size Microaggregation
Jordi Soria-Comas, Josep Domingo-Ferrer, Rafael Mulero-Vellido
MDAI2
2019 Privacy-preserving cloud computing on sensitive data: A survey of methods, products and challenges
Josep Domingo-Ferrer, Oriol Farràs, Jordi Ribes-González, David Sánchez 0001
Comput. Commun.1
2019 Steered Microaggregation as a Unified Primitive to Anonymize Data Sets and Data Streams
abstract
As data grow in quantity and complexity, data anonymization is becoming increasingly challenging. On one side, a great diversity of masking methods, synthetic data generation methods, and privacy models exists, and this diversity is often perceived as unsettling by practitioners. On the other side, most of the anonymization methodology was designed for static, structured, and small data, whereas the current landscape includes big data and, in particular, data streams. We explore here a unified and conceptually simple anonymization approach, by presenting a primitive called steered microaggregation that can be tailored to enforce various privacy models on static data sets and also on data streams. Steered microaggregation is based on adding artificial attributes that are properly initialized and weighted in order to guide the microaggregation process into meeting certain desired constraints. To demonstrate the potential of this type of microaggregation, we show how it can be used to achieve kanonymity, t-closeness, l-diversity, and E-differential privacy in the context of static data sets; furthermore, we discuss how it can be used to achieve k-anonymity of data streams while controlling tuple reordering. Beyond its flexibility and theoretical appeal, steered microaggregation can drastically reduce information loss, as shown by our experimental evaluation.
Josep Domingo-Ferrer, Jordi Soria-Comas, Rafael Mulero-Vellido
IEEE Trans. Inf. Forensics Secur.1
2018 Anonymization of Unstructured Data via Named-Entity Recognition
Fadi Hassan, Josep Domingo-Ferrer, Jordi Soria-Comas
MDAI2
2018 Multiparty Computation with Statistical Input Confidentiality via Randomized Response
Josep Domingo-Ferrer, Rafael Mulero-Vellido, Jordi Soria-Comas
PSD1
2018 On the Privacy Guarantees of Synthetic Data: A Reassessment from the Maximum-Knowledge Attacker Perspective
Nicolás Ruiz, Krishnamurty Muralidhar, Josep Domingo-Ferrer
PSD3
2018 Efficient privacy-preserving implicit authentication
Alberto Blanco-Justicia, Josep Domingo-Ferrer
Comput. Commun.2
2018 Anonymous and secure aggregation scheme in fog-based public cloud computing
Huaqun Wang, Zhiwei Wang 0003, Josep Domingo-Ferrer
Future Gener. Comput. Syst.3
2018 Outsourcing scalar products and matrix products on privacy-protected unencrypted data stored in untrusted clouds
Josep Domingo-Ferrer, Sara Ricci, Carles Domingo-Enrich
Inf. Sci.1
2018 Co-utile disclosure of private data in social networks
David Sánchez 0001, Josep Domingo-Ferrer, Sergio Martínez
Inf. Sci.2
2018 Differentially private data publishing via optimal univariate microaggregation and record perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer
Knowl. Based Syst.2
2017 A Non-Parametric Model for Accurate and Provably Private Synthetic Data Sets
abstract
Generating synthetic data is a well-known option to limit disclosure risk in sensitive data releases. The usual approach is to build a model for the population and then generate a synthetic data set solely based on the model. We argue that building an accurate population model is difficult and we propose instead to approximate the original data as closely as privacy constraints permit. To enforce an ex ante privacy level when generating synthetic data, we introduce a new privacy model called ϵ synthetic privacy. Then, we describe a synthetic data generation method that satisfies ϵ-synthetic privacy. Finally, we evaluate the utility of the synthetic data generated with our method.
Jordi Soria-Comas, Josep Domingo-Ferrer
ARES2
2017 Privacy-Preserving and Co-utile Distributed Social Credit
Josep Domingo-Ferrer
IWOCA1
2017 A Methodology to Compare Anonymization Methods Regarding Their Risk-Utility Trade-off
Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas
MDAI1
2017 Differentially Private Data Sets Based on Microaggregation and Record Perturbation
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI2
2017 Co-Utility: Self-Enforcing protocols for the mutual benefit of participants
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Eng. Appl. Artif. Intell.1
2017 Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees
abstract
Differential privacy is a popular privacy model within the research community because of the strong privacy guarantee it offers, namely that the presence or absence of any individual in a data set does not significantly influence the results of analyses on the data set. However, enforcing this strict guarantee in practice significantly distorts data and/or limits data uses, thus diminishing the analytical utility of the differentially private results. In an attempt to address this shortcoming, several relaxations of differential privacy have been proposed that trade off privacy guarantees for improved data utility. In this paper, we argue that the standard formalization of differential privacy is stricter than required by the intuitive privacy guarantee it seeks. In particular, the standard formalization requires indistinguishability of results between any pair of neighbor data sets, while indistinguishability between the actual data set and its neighbor data sets should be enough. This limits the data controller's ability to adjust the level of protection to the actual data, hence resulting in significant accuracy loss. In this respect, we propose individual differential privacy, an alternative differential privacy notion that offers the same privacy guarantees as standard differential privacy to individuals (even though not to groups of individuals). This new notion allows the data controller to adjust the distortion to the actual data set, which results in less distortion and more analytical accuracy. We propose several mechanisms to attain individual differential privacy and we compare the new notion against standard differential privacy in terms of the accuracy of the analytical results.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, David Megías 0001
IEEE Trans. Inf. Forensics Secur.2
2017 Distributed Aggregate Privacy-Preserving Authentication in VANETs
abstract
Existing secure and privacy-preserving vehicular communication protocols in vehicular ad hoc networks face the challenges of being fast and not depending on ideal tamper-proof devices (TPDs) embedded in vehicles. To address these challenges, we propose a vehicular authentication protocol referred to as distributedaggregate privacy-preserving authentication. The proposed protocol is based on our new multiple trusted authority one-time identity-based aggregate signature technique. With this technique a vehicle can verify many messages simultaneously and their signatures can be compressed into a single one that greatly reduces the storage space needed by a vehicle or a data collector (e.g., the traffic management authority). Instead of ideal TPDs, our protocol only requires realistic TPDs and hence is more practical.
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer, Chuanyan Hu
IEEE Trans. Intell. Transp. Syst.3
2016 t-closeness through microaggregation: Strict privacy with enhanced utility preservation
abstract
This paper proposes and shows how to use microaggregation to attain t-closeness on top of k-anonymity to protect data releases. The advantages in terms of data utility preservation of microaggregation over classic approaches based on generalizing values are analyzed. Then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
ICDE2
2016 Privacy-Preserving Cloud-Based Statistical Analyses on Sensitive Categorical Data
Sara Ricci, Josep Domingo-Ferrer, David Sánchez 0001
MDAI2
2016 Anonymization in the Time of Big Data
Josep Domingo-Ferrer, Jordi Soria-Comas
PSD1
2016 Rank-Based Record Linkage for Re-Identification Risk Assessment
Krishnamurty Muralidhar, Josep Domingo-Ferrer
PSD2
2016 Privacy-aware loyalty programs
Alberto Blanco-Justicia, Josep Domingo-Ferrer
Comput. Commun.2
2016 Big Data Privacy: Challenges to Privacy Principles and Models
abstract
Abstract This paper explores the challenges raised by big data in privacy-preserving data management. First, we examine the conflicts raised by big data with respect to preexisting concepts of private data management, such as consent, purpose limitation, transparency and individual rights of access, rectification and erasure. Anonymization appears as the best tool to mitigate such conflicts, and it is best implemented by adhering to a privacy model with precise privacy guarantees. For this reason, we evaluate how well the two main privacy models used in anonymization (k-anonymity and $$\varepsilon $$ ε -differential privacy) meet the requirements of big data, namely composability, low computational cost and linkability.
Jordi Soria-Comas, Josep Domingo-Ferrer
Data Sci. Eng.2
2016 Cloud Cryptography: Theory, Practice and Future Research Directions
Kim-Kwang Raymond Choo, Josep Domingo-Ferrer, Lei Zhang 0009
Future Gener. Comput. Syst.2
2016 Self-enforcing protocols via co-utile reputation management
Josep Domingo-Ferrer, Oriol Farràs, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Inf. Sci.1
2016 New directions in anonymization: Permutation paradigm, verifiability by subjects and intruders, transparency to users
Josep Domingo-Ferrer, Krishnamurty Muralidhar
Inf. Sci.1
2016 Contributory Broadcast Encryption with Efficient Encryption and Short Ciphertexts
abstract
Broadcast encryption (BE) schemes allow a sender to securely broadcast to any subset of members but require a trusted party to distribute decryption keys. Group key agreement (GKA) protocols enable a group of members to negotiate a common encryption key via open networks so that only the group members can decrypt the ciphertexts encrypted under the shared encryption key, but a sender cannot exclude any particular member from decrypting the ciphertexts. In this paper, we bridge these two notions with a hybrid primitive referred to as contributory broadcast encryption (ConBE). In this new primitive, a group of members negotiate a common public encryption key while each member holds a decryption key. A sender seeing the public group encryption key can limit the decryption to a subset of members of his choice. Following this model, we propose a ConBE scheme with short ciphertexts. The scheme is proven to be fully collusion-resistant under the decision n-Bilinear Diffie-Hellman Exponentiation (BDHE) assumption in the standard model. Of independent interest, we present a new BE scheme that is aggregatable. The aggregatability property is shown to be useful to construct advanced protocols.
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer, Oriol Farràs, Jesús A. Manjón
IEEE Trans. Computers4
2016 Privacy-Preserving Vehicular Communication Authentication with Hierarchical Aggregation and Fast Response
abstract
Existing secure and privacy-preserving schemes for vehicular communications in vehicular ad hoc networks face some challenges, e.g., reducing the dependence on ideal tamper-proof devices, building efficient member revocation mechanisms and avoiding computation and communication bottlenecks. To cope with those challenges, we propose a highly efficient secure and privacy-preserving scheme based on identity-based aggregate signatures. Our scheme enables hierarchical aggregation and batch verification. The individual identity-based signatures generated by different vehicles can be aggregated and verified in a batch. The aggregated signatures can be re-aggregated by a message collector (e.g., traffic management authority). With our hierarchical aggregation technique, we significantly reduce the transmission/storage overhead of the vehicles and other parties. Furthermore, existing batch verification based schemes in vehicular ad hoc networks require vehicles to wait for enough messages to perform a batch verification. In contrast, we assume that vehicles will generate messages (and the corresponding signatures) in certain time spans, so that vehicles only need to wait for a very short period before they can start the batch verification procedure. Simulation shows that a vehicle can verify the received messages with very low latency and fast response.
Lei Zhang 0009, Chuanyan Hu, Qianhong Wu, Josep Domingo-Ferrer
IEEE Trans. Computers4
2015 Co-utile Collaborative Anonymization of Microdata
Jordi Soria-Comas, Josep Domingo-Ferrer
MDAI2
2015 Disclosure risk assessment via record linkage by a maximum-knowledge attacker
abstract
Before releasing an anonymized data set, the data protector must know how safe the data set is, that is, how much disclosure risk is incurred by the release. If no privacy model is used to select specific privacy guarantees prior to anonymization, posterior disclosure risk assessment must be performed based on the anonymized data set and, if the result is not satisfactory, anonymization must be repeated with stricter privacy parameters. Even if a privacy model is used, it may still be advisable to empirically evaluate disclosure on the anonymized data set, especially if the privacy model parameters have been relaxed to improve data utility. Record linkage is a general methodology to posterior disclosure risk assessment, whereby the data protector attempts to recreate the attacker's re-identification scenario. An important limitation of record linkage is that it usually requires the data protector to make restrictive assumptions on the attacker's background knowledge. To overcome this limitation, we present a maximum-knowledge attacker model and then we specify and compare several record linkage tests for such a worst-case attacker. Our tests are based on comparing the distribution of linkage distances between the original and the anonymized data set with the distribution of distances between one of the two previous data sets and one random data set. The more similar the distributions, the more plausibly deniable are record linkages claimed by an attacker. Because attaining zero disclosure risk for all records is too costly in terms of utility, a less demanding alternative is presented whose goal is to reduce the maximum per-record disclosure risk.
Josep Domingo-Ferrer, Sara Ricci, Jordi Soria-Comas
PST1
2015 Flexible and Robust Privacy-Preserving Implicit Authentication
Josep Domingo-Ferrer, Qianhong Wu, Alberto Blanco-Justicia
SEC1
2015 Practical secure and privacy-preserving scheme for value-added applications in VANETs
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
Comput. Commun.4
2015 Security and privacy in unified communications: Challenges and solutions
Georgios Karopoulos, Georgios Portokalidis, Josep Domingo-Ferrer, Ying-Dar Lin, Dimitris Geneiatakis, Georgios Kambourakis
Comput. Commun.3
2015 Discrimination- and privacy-aware patterns
Sara Hajian, Josep Domingo-Ferrer, Anna Monreale, Dino Pedreschi, Fosca Giannotti
Data Min. Knowl. Discov.2
2015 Semantic variance: An intuitive measure for ontology accuracy evaluation
David Sánchez 0001, Montserrat Batet, Sergio Martínez, Josep Domingo-Ferrer
Eng. Appl. Artif. Intell.4
2015 A k-anonymous approach to privacy preserving collaborative filtering
Fran Casino, Josep Domingo-Ferrer, Constantinos Patsakis, Domenec Puig, Agusti Solanas
J. Comput. Syst. Sci.2
2015 From t-closeness to differential privacy and vice versa in data anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas
Knowl. Based Syst.1
2015 TPP: Traceable Privacy-Preserving Communication and Precise Reward for Vehicle-to-Grid Networks in Smart Grids
abstract
In vehicle-to-grid (V2G) networks, service providers are battery-powered vehicles, and the service consumer is the power grid. Security and privacy concerns are major obstacles for V2G networks to be extensively deployed. In 2011, Yang et al. proposed a very interesting privacy-preserving communication and precise reward architecture for V2G networks in smart grids. In this paper, we enhance Yang et al.'s framework with the formal definitions of unforgeability and restrictiveness. Then, we propose a new traceable privacy-preserving communication and precise reward scheme with available cryptographic primitives. The proposed scheme is formally proven secure with well-established assumptions in the random oracle model. Thorough theoretical and experimental analyses demonstrate that our scheme is efficient and practical for secure V2G networks in smart grids.
Huaqun Wang, Qianhong Wu, Li Xu 0002, Josep Domingo-Ferrer
IEEE Trans. Inf. Forensics Secur.5
2015 Generating Searchable Public-Key Ciphertexts With Hidden Structures for Fast Keyword Search
abstract
Existing semantically secure public-key searchable encryption schemes take search time linear with the total number of the ciphertexts. This makes retrieval from large-scale databases prohibitive. To alleviate this problem, this paper proposes searchable public-key ciphertexts with hidden structures (SPCHS) for keyword search as fast as possible without sacrificing semantic security of the encrypted keywords. In SPCHS, all keyword-searchable ciphertexts are structured by hidden relations, and with the search trapdoor corresponding to a keyword, the minimum information of the relations is disclosed to a search algorithm as the guidance to find all matching ciphertexts efficiently. We construct an SPCHS scheme from scratch in which the ciphertexts have a hidden star-like structure. We prove our scheme to be semantically secure in the random oracle (RO) model. The search complexity of our scheme is dependent on the actual number of the ciphertexts containing the queried keyword, rather than the number of all ciphertexts. Finally, we present a generic SPCHS construction from anonymous identity-based encryption and collision-free full-identity malleable identity-based key encapsulation mechanism (IBKEM) with anonymity. We illustrate two collision-free full-identity malleable IBKEM instances, which are semantically secure and anonymous, respectively, in the RO and standard models. The latter instance enables us to construct an SPCHS scheme with semantic security in the standard model.
Peng Xu 0003, Qianhong Wu, Wei Wang 0088, Willy Susilo, Josep Domingo-Ferrer, Hai Jin 0001
IEEE Trans. Inf. Forensics Secur.5
2015 Round-Efficient and Sender-Unrestricted Dynamic Group Key Agreement Protocol for Secure Group Communications
abstract
Modern collaborative and group-oriented applications typically involve communications over open networks. Given the openness of today's networks, communications among group members must be secure and, at the same time, efficient. Group key agreement (GKA) is widely employed for secure group communications in modern collaborative and group-oriented applications. This paper studies the problem of GKA in identity-based cryptosystems with an emphasis on round-efficient, sender-unrestricted, member-dynamic, and provably secure key escrow freeness. The problem is resolved by proposing a one-round dynamic asymmetric GKA protocol which allows a group of members to dynamically establish a public group encryption key, while each member has a different secret decryption key in an identity-based cryptosystem. Knowing the group encryption key, any entity can encrypt to the group members so that only the members can decrypt. We construct this protocol with a strongly unforgeable stateful identity-based batch multisignature scheme. The proposed protocol is shown to be secure under the k -bilinear Diffie-Hellman exponent assumption.
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer, Zheming Dong
IEEE Trans. Inf. Forensics Secur.3
2015 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation
abstract
Microaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate k-anonymous data sets, where the identity of each subject is hidden within a group of k subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers, and avoiding discretization of numerical data. k-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of k subjects is too small. To address this issue, several refinements of k-anonymity have been proposed, among which t-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate t-close data sets are based on generalization and suppression (they are extensions of k-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate k-anonymous t-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
IEEE Trans. Knowl. Data Eng.2
2014 Tracing and revoking leaked credentials: accountability in leaking sensitive outsourced data
abstract
Most existing proposals for access control over outsourced data mainly aim at guaranteeing that the data are only accessible to authorized requestors who have the access credentials. This paper proposes TRLAC, an a posteriori approach for tracing and revoking leaked credentials, to complement existing a priori solutions. The tracing procedure of TRLAC can trace, in a black-box manner, at least one traitor who illegally distributed a credential, without any help from the cloud service provider. Once the dishonest users have been found, a revocation mechanism can be called to deprive them of access rights. We formally prove the security of TRLAC, and empirically shows that the introduction of the tracing feature incurs little costs to outsourcing.
Qianhong Wu, Sherman S. M. Chow, Josep Domingo-Ferrer, Wenchang Shi
AsiaCCS5
2014 Data Anonymization
Josep Domingo-Ferrer, Jordi Soria-Comas
CRiSIS1
2014 A Provably Secure Ring Signature Scheme with Bounded Leakage Resilience
Huaqun Wang, Qianhong Wu, Futai Zhang, Josep Domingo-Ferrer
ISPEC5
2014 Improving the Utility of Differential Privacy via Univariate Microaggregation
David Sánchez 0001, Josep Domingo-Ferrer, Sergio Martínez
Privacy in Statistical Databases2
2014 Reverse Mapping to Preserve the Marginal Distributions of Attributes in Masked Microdata
Krishnamurty Muralidhar, Rathindra Sarathy, Josep Domingo-Ferrer
Privacy in Statistical Databases3
2014 Distance Computation between Two Private Preference Functions
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Oriol Farràs, David Sánchez 0001
SEC2
2014 Security in wireless ad-hoc networks - A survey
Roberto Di Pietro, Stefano Guarino, Nino Vincenzo Verde, Josep Domingo-Ferrer
Comput. Commun.4
2014 Generalization-based privacy preservation and discrimination prevention in data publishing and mining
Sara Hajian, Josep Domingo-Ferrer, Oriol Farràs
Data Min. Knowl. Discov.2
2014 Identity-based remote data possession checking in public clouds
abstract
Checking remote data possession is of crucial importance in public cloud storage. It enables the users to check whether their outsourced data have been kept intact without downloading the original data. The existing remote data possession checking (RDPC) protocols have been designed in the PKI (public key infrastructure) setting. The cloud server has to validate the users’ certificates before storing the data uploaded by the users in order to prevent spam. This incurs considerable costs since numerous users may frequently upload data to the cloud server. This study addresses this problem with a new model of identity‐based RDPC (ID‐RDPC) protocols. The authors present the first ID‐RDPC protocol proven to be secure assuming the hardness of the standard computational Diffie‐Hellman problem. In addition to the structural advantage of elimination of certificate management and verification, the authors ID‐RDPC protocol also outperforms the existing RDPC protocols in the PKI setting in terms of computation and communication.
Huaqun Wang, Qianhong Wu, Josep Domingo-Ferrer
IET Inf. Secur.4
2014 Ciphertext-policy hierarchical attribute-based encryption with short ciphertexts
Qianhong Wu, Josep Domingo-Ferrer, Lei Zhang 0009, Jianwei Liu 0001, Wenchang Shi
Inf. Sci.4
2014 Signatures in hierarchical certificateless cryptography: Efficient constructions and provable security
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.3
2014 FRR: Fair remote retrieval of outsourced private medical records in electronic health networks
Huaqun Wang, Qianhong Wu, Josep Domingo-Ferrer
J. Biomed. Informatics4
2014 Privacy-aware peer-to-peer content distribution using automatically recombined fingerprints
David Megías 0001, Josep Domingo-Ferrer
Multim. Syst.2
2014 Enhancing data utility in differential privacy via microaggregation-based k-anonymity
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
VLDB J.2
2013 A Generic Construction of Proxy Signatures from Certificateless Signatures
abstract
The primitive of proxy signatures allows the original signer to delegate proxy signers to sign on messages on behalf of the original signer. It has found numerous applications in distributed computing scenarios where delegation of signing rights is common. Certificate less public key cryptography eliminates the complicated certificates in traditional public key cryptosystems without suffering from the key escrow problem in identity-based public key cryptography. In this paper, we reveal the relationship between the two important primitives of proxy signatures and certificate less signatures and present a generic conversion from the latter to the former. Following the generic transformation, we propose an efficient proxy signature scheme with a recent certificate less signature scheme.
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer, Jianwei Liu 0001, Ruiying Du
AINA4
2013 DNA-inspired anonymous fingerprinting for efficient peer-to-peer content distribution
abstract
When selling electronic content, the merchant would like each buyer to receive a different copy of the content fingerprinted with a serial number, in order to be able to trace redistributors should illegal redistribution happen. On the other hand, the merchant would like content distribution to be as scalable as possible, in order for mass transactions to be possible. Multicast content distribution fails to satisfy the first requirement: all receivers get exactly the same copy of the content, which makes it difficult to trace illegal redistributors. Unicast distribution of fingerprinted content, on the other hand, fails to satisfy the second requirement: for each buyer, the merchant needs to compute a fingerprint and establish a connection. P2P content distribution is a third option combining the strengths of multicast and unicast: the merchant needs to establish unicast connections only with a few seed buyers; on the other hand, with a suitable fingerprinting mechanism, illegal redistributors can still be identified and honest buyers can stay anonymous. We present a P2P content distribution scheme with such an anonymous fingerprinting mechanism, which is inspired in the way DNA sequences combine and spread from ancestors to descendants.
David Megías 0001, Josep Domingo-Ferrer
IEEE Congress on Evolutionary Computation2
2013 Secure One-to-Group Communications Escrow-Free ID-Based Asymmetric Group Key Agreement
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer, Sherman S. M. Chow, Wenchang Shi
Inscrypt3
2013 Facility Location and Social Choice via Microaggregation
Josep Domingo-Ferrer
MDAI1
2013 Differential privacy via t-closeness in data publishing
abstract
k-Anonymity and e-differential privacy are two main privacy models proposed within the computer science community. Whereas the former was proposed for privacy-preserving data publishing, i.e. data set anonymization, the latter initially arose in the context of interactive databases and was later extended to data publishing. We show here that t-closeness, one of the extensions of k-anonymity, can actually yieldε-differential privacy in data publishing when t =exp(ε). We detail a construction based on bucketization that realizes the previous implication; hence, as an ancillary result, we provide a new computational procedure to achieve t-closeness and ε-differential privacy in data publishing.
Jordi Soria-Comas, Josep Domingo-Ferrer
PST2
2013 On the Connection between t-Closeness and Differential Privacy for Data Releases
Josep Domingo-Ferrer
SECRYPT1
2013 Distributed multicast of fingerprinted content based on a rational peer-to-peer community
Josep Domingo-Ferrer, David Megías 0001
Comput. Commun.1
2013 On the privacy offered by (k, δ)-anonymity
Rolando Trujillo-Rasua, Josep Domingo-Ferrer
Inf. Syst.2
2013 Anonymization of nominal data based on semantic marginality
Josep Domingo-Ferrer, David Sánchez 0001, Guillem Rufian-Torrell
Inf. Sci.1
2013 Optimal data-independent noise for differential privacy
Jordi Soria-Comas, Josep Domingo-Ferrer
Inf. Sci.2
2013 A Methodology for Direct and Indirect Discrimination Prevention in Data Mining
abstract
Data mining is an increasingly important technology for extracting useful knowledge hidden in large collections of data. There are, however, negative social perceptions about data mining, among which potential privacy invasion and potential discrimination. The latter consists of unfairly treating people on the basis of their belonging to a specific group. Automated data collection and data mining techniques such as classification rule mining have paved the way to making automated decisions, like loan granting/denial, insurance premium computation, etc. If the training data sets are biased in what regards discriminatory (sensitive) attributes like gender, race, religion, etc., discriminatory decisions may ensue. For this reason, anti-discrimination techniques including discrimination discovery and prevention have been introduced in data mining. Discrimination can be either direct or indirect. Direct discrimination occurs when decisions are made based on sensitive attributes. Indirect discrimination occurs when decisions are made based on nonsensitive attributes which are strongly correlated with biased sensitive ones. In this paper, we tackle discrimination prevention in data mining and propose new techniques applicable for direct or indirect discrimination prevention individually or both at the same time. We discuss how to clean training data sets and outsourced data sets in such a way that direct and/or indirect discriminatory decision rules are converted to legitimate (nondiscriminatory) classification rules. We also propose new metrics to evaluate the utility of the proposed approaches and we compare these approaches. The experimental evaluations demonstrate that the proposed techniques are effective at removing direct and/or indirect discrimination biases in the original data set while preserving data quality.
Sara Hajian, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.2
2013 Fast Transmission to Remote Cooperative Groups: A New Key Management Paradigm
abstract
The problem of efficiently and securely broadcasting to a remote cooperative group occurs in many newly emerging networks. A major challenge in devising such systems is to overcome the obstacles of the potentially limited communication from the group to the sender, the unavailability of a fully trusted key generation center, and the dynamics of the sender. The existing key management paradigms cannot deal with these challenges effectively. In this paper, we circumvent these obstacles and close this gap by proposing a novel key management paradigm. The new paradigm is a hybrid of traditional broadcast encryption and group key agreement. In such a system, each member maintains a single public/secret key pair. Upon seeing the public keys of the members, a remote sender can securely broadcast to any intended subgroup chosen in an ad hoc way. Following this model, we instantiate a scheme that is proven secure in the standard model. Even if all the nonintended members collude, they cannot extract any useful information from the transmitted messages. After the public group encryption key is extracted, both the computation overhead and the communication cost are independent of the group size. Furthermore, our scheme facilitates simple yet efficient member deletion/addition and flexible rekeying strategies. Its strong security against collusion, its constant overhead, and its implementation friendliness without relying on a fully trusted authority render our protocol a very promising solution to many applications.
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer, Jesús A. Manjón
IEEE/ACM Trans. Netw.4
2012 Probabilistic k-anonymity through microaggregation and data swapping
abstract
k-Anonymity is a privacy property used to limit the risk of re-identification in a microdata set. A data set satisfying k-anonymity consists of groups of k records which are indistinguishable as far as their quasi-identifier attributes are concerned. Hence, the probability of re-identifying a record within a group is 1/k. We introduce the probabilistic k-anonymity property, which relaxes the indistinguishability requirement of k-anonymity and only requires that the probability of re-identification be the same as in k-anonymity. Two computational heuristics to achieve probabilistic k-anonymity based on data swapping are proposed: MDAV microaggregation on the quasi-identifiers plus swapping, and individual ranking microaggregation on individual confidential attributes plus swapping. We report experimental results, where we compare the utility of original, k-anonymous and probabilistically k-anonymous data.
Jordi Soria-Comas, Josep Domingo-Ferrer
FUZZ-IEEE2
2012 Marginality: A Numerical Mapping for Enhanced Exploitation of Taxonomic Attributes
Josep Domingo-Ferrer
MDAI1
2012 Anonymization Methods for Taxonomic Microdata
Josep Domingo-Ferrer, Krishnamurty Muralidhar, Guillem Rufian-Torrell
Privacy in Statistical Databases1
2012 Hybrid Microdata via Model-Based Clustering
Anna Oganian, Josep Domingo-Ferrer
Privacy in Statistical Databases2
2012 Sensitivity-Independent differential Privacy via Prior Knowledge Refinement
abstract
We propose a new mechanism to implement differential privacy. Unlike the usual mechanism based on adding a noise whose magnitude is proportional to the sensitivity of the query function, our proposal is based on the refinement of the user's prior knowledge about the response. Our mechanism is shown to have several advantages over noise addition: it does not require complex computations, and thus it can be easily automated; it lets the user exploit her prior knowledge about the response to achieve better data quality; and it is independent of the sensitivity of the query function (although this can be a disadvantage if the sensitivity is small). Furthermore, we give a general algorithm for knowledge refinement and we show some compounding properties of our mechanism for the case of multiple queries; also, we build an interactive mechanism on top of knowledge refinement and we show that it is safe against adaptive attacks. Finally, we give a quality assessment for the responses to individual queries.
Jordi Soria-Comas, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2012 Rational behavior in peer-to-peer profile obfuscation for anonymous keyword search
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2012 Rational behavior in peer-to-peer profile obfuscation for anonymous keyword search: The multi-hop scenario
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2012 Microaggregation- and permutation-based anonymization of movement data
Josep Domingo-Ferrer, Rolando Trujillo-Rasua
Inf. Sci.1
2012 Provably secure threshold public-key encryption with adaptive security and short ciphertexts
Qianhong Wu, Lei Zhang 0009, Oriol Farràs, Josep Domingo-Ferrer
Inf. Sci.5
2012 Query Profile Obfuscation by Means of Optimal Query Exchange between Users
abstract
We address the problem of query profile obfuscation by means of partial query exchanges between two users, in order for their profiles of interest to appear distorted to the information provider (database, search engine, etc.). We illustrate a methodology to reach mutual privacy gain, that is, a situation where both users increase their own privacy protection through collaboration in query exchange. To this end, our approach starts with a mathematical formulation, involving the modeling of the users' apparent profiles as probability distributions over categories of interest, and the measure of their privacy as the corresponding Shannon entropy. The question of which query categories to exchange translates into finding optimization variables representing exchange policies, for various optimization objectives based on those entropies, possibly under exchange traffic constraints.
David Rebollo-Monedero, Jordi Forné, Josep Domingo-Ferrer
IEEE Trans. Dependable Secur. Comput.3
2011 Bridging Broadcast Encryption and Group Key Agreement
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer, Oriol Farràs
ASIACRYPT4
2011 Discrimination prevention in data mining for intrusion and crime detection
abstract
Automated data collection has fostered the use of data mining for intrusion and crime detection. Indeed, banks, large corporations, insurance companies, casinos, etc. are increasingly mining data about their customers or employees in view of detecting potential intrusion, fraud or even crime. Mining algorithms are trained from datasets which may be biased in what regards gender, race, religion or other attributes. Furthermore, mining is often outsourced or carried out in cooperation by several entities. For those reasons, discrimination concerns arise. Potential intrusion, fraud or crime should be inferred from objective misbehavior, rather than from sensitive attributes like gender, race or religion. This paper discusses how to clean training datasets and outsourced datasets in such a way that legitimate classification rules can still be extracted but discriminating rules based on sensitive attributes cannot.
Sara Hajian, Josep Domingo-Ferrer, Antoni Martínez-Ballesté
CICS2
2011 Preserving Security and Privacy in Large-Scale VANETs
Qianhong Wu, Josep Domingo-Ferrer, Lei Zhang 0009
ICICS3
2011 APPA: Aggregate Privacy-Preserving Authentication in Vehicular Ad Hoc Networks
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
ISC4
2011 Rule Protection for Indirect Discrimination Prevention in Data Mining
Sara Hajian, Josep Domingo-Ferrer, Antoni Martínez-Ballesté
MDAI2
2011 Fully Distributed Broadcast Encryption
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer
ProvSec4
2011 Decapitation of networks with and without weights and direction: The economics of iterated attack and defense
Josep Domingo-Ferrer, Úrsula González-Nicolás
Comput. Networks1
2011 Asymmetric group key agreement protocol for open networks and its application to broadcast encryption
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer, Úrsula González-Nicolás
Comput. Networks4
2011 Co-citations and Relevance of Authors and Author Groups
abstract
The way an author or a group of authors are cited tells more about the real impact of their work than authorship and collaborations. Indeed, the connections within the scientific community can be more accurately elicited from the co-citation graph than from the collaboration graph. We suggest some indices that can be drawn from the co-citation graph in order to capture the relevance of individual authors and the relevance of groups of authors.
Maria Bras-Amorós, Josep Domingo-Ferrer, Albert Vico-Oton
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2011 Provably secure one-round identity-based authenticated asymmetric group key agreement protocol
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.4
2010 Ad hoc broadcast encryption
abstract
Numerous applications in ad hoc networks, peer-to-peer networks, and on-the-fly data sharing call for confidential broadcast without relying on a dealer. To cater for such applications, we propose a new primitive referred to as ad hoc broadcast encryption (AHBE), in which each user possesses a public key and, upon seeing the public keys of the users, a sender can securely broadcast to any subset of them, so that only the intended users can decrypt. We implement a concrete AHBE scheme proven secure under the decision Bilinear Diffie-Hellman Exponentiation (BDHE) assumption. The resulting scheme has sub-linear complexity, comparable to up-to-date broadcast systems which have also sub-linear complexity but require a fully trusted dealer.
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer
CCS4
2010 Identity-Based Authenticated Asymmetric Group Key Agreement Protocol
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
COCOON4
2010 Hierarchical Certificateless Signatures
abstract
Certificateless cryptography eliminates the key escrow problem in identity-based cryptography. Hierarchical cryptography exploits a practical security model to mirror the organizational hierarchy in the real world. In this paper, to incorporate the advantages of both types of cryptosystems, we instantiate hierarchical certificate less cryptography by formalizing the notion of hierarchical certificate less signatures. Furthermore, we propose an HCLS scheme which, under the hardness of the computational Diffie-Hellman (CDH) problem, is proven to be existentially unforgeable against adaptive chosen-message attacks in the random oracle model. As to efficiency, our scheme has constant complexity, regardless of the depth of the hierarchy. Hence, the proposal is secure and scalable for practical applications.
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
EUC3
2010 Threshold Public-Key Encryption with Adaptive Security and Short Ciphertexts
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer
ICICS4
2010 A Bibliometric Index Based on Collaboration Distances
Maria Bras-Amorós, Josep Domingo-Ferrer, Vicenç Torra
MDAI2
2010 User Privacy in Web Search
Josep Domingo-Ferrer
MDAI1
2010 Rational Privacy Disclosure in Social Networks
Josep Domingo-Ferrer
MDAI1
2010 Coprivacy: Towards a Theory of Sustainable Privacy
Josep Domingo-Ferrer
Privacy in Statistical Databases1
2010 Secure compression of privacy-preserving witnesses in vehicular ad hoc networks
abstract
Vehicular ad hoc networks (VANETs) are designed to improve traffic safety and efficiency. To this end, the traffic communication must be authenticated to guarantee trustworthiness for guiding drivers and establishing liability in case of traffic accident investigation. Cryptographic authentication techniques have been extensively exploited to secure VANETs. Applying cryptographic authentication techniques such as digital signatures raises challenges to efficiently store signatures on messages growing with time. To alleviate the conflict between traffic liability investigation and limited storage capacity in vehicles, this paper proposes to aggregate signatures in VANETs. Our proposal can preserve privacy for honest vehicles and trace misbehaving ones, and provides a practical balance between security and privacy in VANETs. With our proposal, cryptographic witnesses of safety-related traffic messages can be significantly compressed so that they can be stored for a long period for liability investigation. Our proposal allows a large number of traffic messages to be verified as if they were a single one, which speeds up the response of vehicles to traffic reports.
Qianhong Wu, Lei Zhang 0009, Josep Domingo-Ferrer
WiMob4
2010 Hybrid microdata using microaggregation
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2010 Simulatable certificateless two-party authenticated key agreement protocol
Lei Zhang 0009, Futai Zhang, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.4
2010 From t-Closeness-Like Privacy to Postrandomization via Information Theory
abstract
t-Closeness is a privacy model recently defined for data anonymization. A data set is said to satisfy t-closeness if, for each group of records sharing a combination of key attributes, the distance between the distribution of a confidential attribute in the group and the distribution of the attribute in the entire data set is no more than a threshold t. Here, we define a privacy measure in terms of information theory, similar to t-closeness. Then, we use the tools of that theory to show that our privacy measure can be achieved by the postrandomization method (PRAM) for masking in the discrete case, and by a form of noise addition in the general case.
David Rebollo-Monedero, Jordi Forné, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.3
2009 Asymmetric Group Key Agreement
Qianhong Wu, Yi Mu 0001, Willy Susilo, Josep Domingo-Ferrer
EUROCRYPT5
2009 The Functionality-Security-Privacy Game
Josep Domingo-Ferrer
MDAI1
2009 Aggregation of Trustworthy Announcement Messages in Vehicular Ad Hoc Networks
abstract
Vehicular ad hoc networks (VANETs) allow vehicle-to-vehicle communication and, in particular, vehicle-generated announcements. Vehicles can use such announcements to warn nearby vehicles about road conditions (traffic jams, accidents). Thus, they can greatly increase the safety of driving. However, their trustworthiness must be guaranteed. A new system for vehicle-generated announcements is presented that is secure against external and internal attackers attempting to send fake messages. Internal attacks are thwarted by using an endorsement mechanism based on multisignatures. Besides, this scheme ensures that vehicles volunteering to generate and/or endorse trustworthy announcements do not have to sacrifice their privacy.
Alexandre Viejo, Francesc Sebé, Josep Domingo-Ferrer
VTC Spring3
2009 User-private information retrieval based on a peer-to-peer community
Josep Domingo-Ferrer, Maria Bras-Amorós, Qianhong Wu, Jesús A. Manjón
Data Knowl. Eng.1
2009 Recent progress in database privacy
Josep Domingo-Ferrer, Yücel Saygin
Data Knowl. Eng.1
2009 Erratum to "A measure of variance for hierarchical nominal attributes"
Josep Domingo-Ferrer, Agusti Solanas
Inf. Sci.1
2008 A Critique of k-Anonymity and Some of Its Enhancements
abstract
k-Anonymity is a privacy property requiring that all combinations of key attributes in a database be repeated at least for k records. It has been shown that k-anonymity alone does not always ensure privacy. A number of sophistications of k-anonymity have been proposed, like p-sensitive k-anonymity, l-diversity and t-closeness. This paper explores the shortcomings of those properties, none of which turns out to be completely convincing.
Josep Domingo-Ferrer, Vicenç Torra
ARES1
2008 On intuitionistic fuzzy clustering for its application to privacy
abstract
Motivated by our research on specific information loss measures (in privacy preserving data mining) and our need to compare fuzzy clusters, we proposed in a recent paper a definition for intuitionistic fuzzy partitions. We showed how to define them in the framework of fuzzy clustering. That is, we introduced a method to define intuitionistic fuzzy partitions from the results of fuzzy clustering. In this paper we further study such intuitionistic fuzzy partitions and we extend our previous results with other types of fuzzy clustering algorithms.
Vicenç Torra, Sadaaki Miyamoto, Yasunori Endo, Josep Domingo-Ferrer
FUZZ-IEEE4
2008 A Shared Steganographic File System with Error Correction
Josep Domingo-Ferrer, Maria Bras-Amorós
MDAI1
2008 Peer-to-Peer Private Information Retrieval
Josep Domingo-Ferrer, Maria Bras-Amorós
Privacy in Statistical Databases1
2008 From t-Closeness to PRAM and Noise Addition Via Information Theory
David Rebollo-Monedero, Jordi Forné, Josep Domingo-Ferrer
Privacy in Statistical Databases3
2008 Privacy homomorphisms for social networks with private relationships
Josep Domingo-Ferrer, Alexandre Viejo, Francesc Sebé, Úrsula González-Nicolás
Comput. Networks1
2008 Secure and scalable many-to-one symbol transmission for sensor networks
Alexandre Viejo, Francesc Sebé, Josep Domingo-Ferrer
Comput. Commun.3
2008 A measure of variance for hierarchical nominal attributes
Josep Domingo-Ferrer, Agusti Solanas
Inf. Sci.1
2008 Efficient Remote Data Possession Checking in Critical Information Infrastructures
abstract
Checking data possession in networked information systems such as those related to critical infrastructures (power facilities, airports, data vaults, defense systems, etc.) is a matter of crucial importance. Remote data possession checking protocols permit to check that a remote server can access an uncorrupted file in such a way that the verifier does not need to know beforehand the entire file that is being verified. Unfortunately, current protocols only allow a limited number of successive verifications or are impractical from the computational point of view. In this paper, we present a new remote data possession checking protocol such that: i) it allows an unlimited number of file integrity verifications; ii) its maximum running time can be chosen at set-up time and traded off against storage at the verifier.
Francesc Sebé, Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Yves Deswarte, Jean-Jacques Quisquater
IEEE Trans. Knowl. Data Eng.2
2007 An Incentive-Based System for Information Providers over Peer-to-Peer Mobile Ad-Hoc Networks
Jordi Castellà-Roca, Vanesa Daza, Josep Domingo-Ferrer, Jesús A. Manjón, Francesc Sebé, Alexandre Viejo
MDAI3
2007 A Public-Key Protocol for Social Networks with Private Relationships
Josep Domingo-Ferrer
MDAI1
2007 Advances in smart cards
Josep Domingo-Ferrer, Joachim Posegga, Francesc Sebé, Vicenç Torra
Comput. Networks1
2007 Scalability and security in biased many-to-one communication
Francesc Sebé, Josep Domingo-Ferrer
Comput. Networks2
2007 Secure many-to-one symbol transmission for implementation on smart cards
Francesc Sebé, Alexandre Viejo, Josep Domingo-Ferrer
Comput. Networks3
2007 A distributed architecture for scalable private RFID tag identification
Agusti Solanas, Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Vanesa Daza
Comput. Networks2
2006 A 2d-Tree-Based Blocking Method for Microaggregating Very Large Data Sets
abstract
Blocking is a well-known technique used to partition a set of records into several subsets of manageable size. The standard approach to blocking is to split the records according to the values of one or several attributes (called blocking attributes). This paper presents a new blocking method based on 2/sup d/-trees for intelligently partitioning very large data sets for micro aggregation. A number of experiments has been carried out in order to compare our method with the most typical univariate one.
Agusti Solanas, Antoni Martínez-Ballesté, Josep Domingo-Ferrer, Josep Maria Mateo-Sanz
ARES3
2006 A Smart Card-Based Mental Poker System
Jordi Castellà-Roca, Josep Domingo-Ferrer, Francesc Sebé
CARDIS2
2006 Establishing a benchmark for re-identification methods and its validation using fuzzy clustering
abstract
Privacy preserving data mining and statistical disclosure control are related fields with increasing importance nowadays. They aim is to allow the publication of sensible data without compromising the privacy of data respondents. To that end, masking methods have been designed so that data are distorted in a way that preserves confidentiality and data utility. Alternatively, methods have been constructed to generate synthetic data that have properties similar to the ones of the original data. At the same time, recent research in re-identification methods (record and variable matching) has been pushed forward due to the current interest on security issues and the huge amount of data stored in databases. However, there is no standard methodology for comparing alternative re-identification methods. In this paper we propose the use of masking methods and synthetic data generators for building benchmarks for matching methods. We validate our approach using fuzzy clustering.
Vicenç Torra, Josep Domingo-Ferrer
FUZZ-IEEE2
2006 Watermarking Non-numerical Databases
Agusti Solanas, Josep Domingo-Ferrer
MDAI2
2006 Optimal Multivariate 2-Microaggregation for Microdata Protection: A 2-Approximation
Josep Domingo-Ferrer, Francesc Sebé
Privacy in Statistical Databases1
2006 Using Mahalanobis Distance-Based Record Linkage for Disclosure Risk Assessment
Vicenç Torra, John M. Abowd, Josep Domingo-Ferrer
Privacy in Statistical Databases3
2006 Watermarking Numerical Data in the Presence of Noise
abstract
Data mining aims at extracting knowledge from data. Sometimes this requires the data owner to make data available to the data analyst. Unless the analyst is trusted by the data owner, the latter may wish to have his intellectual property rights protected. Watermarking is a way to provide some protection, not only for multimedia data, but also for numerical data. It consists of imperceptibly embedding a secret mark into the data to be protected. The mark can later be used to resolve any subsequent disputes on data ownership. This paper presents a watermarking method that embeds a watermark into each attribute of a multivariate continuous numerical dataset. This watermarking method is shown to be robust against random noise addition attacks. Data quality is assured to the extent that the watermarked data nearly preserve the attribute means and the covariance matrix from the original dataset. Our proposal is the first known watermarking system for multivariate numerical datasets with such robustness and quality properties.
Francesc Sebé, Josep Domingo-Ferrer, Jordi Castellà-Roca
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2006 Regression for ordinal variables without underlying continuous variables
Vicenç Torra, Josep Domingo-Ferrer, Josep Maria Mateo-Sanz, Michael Kwok-Po Ng
Inf. Sci.2
2006 Efficient multivariate data-oriented microaggregation
Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Josep Maria Mateo-Sanz, Francesc Sebé
VLDB J.1
2005 Noise-Robust Watermarking for Numerical Datasets
Francesc Sebé, Josep Domingo-Ferrer, Agusti Solanas
MDAI2
2005 Dropout-Tolerant TTP-Free Mental Poker
Jordi Castellà-Roca, Francesc Sebé, Josep Domingo-Ferrer
TrustBus3
2005 Privacy in Data Mining
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.1
2005 Ordinal, Continuous and Heterogeneous k-Anonymity Through Microaggregation
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.1
2005 Probabilistic Information Loss Measures in Confidentiality Protection of Continuous Microdata
Josep Maria Mateo-Sanz, Josep Domingo-Ferrer, Francesc Sebé
Data Min. Knowl. Discov.2
2004 Object Positioning Based on Partial Preferences
Josep Maria Mateo-Sanz, Josep Domingo-Ferrer, Vicenç Torra
MDAI2
2004 Secure Reverse Communication in a Multicast Tree
Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Francesc Sebé
NETWORKING1
2004 Large-Scale Pay-As-You-Watch for Unicast and Multicast Communications
Antoni Martínez-Ballesté, Francesc Sebé, Josep Domingo-Ferrer
TrustBus3
2003 Median-based aggregation operators for prototype construction in ordinal scales
abstract
This article studies aggregation operators in ordinal scales for their application to clustering (more specifically, to microaggregation for statistical disclosure risk). In particular, we consider these operators in the process of prototype construction. This study analyzes main aggregation operators for ordinal scales [plurality rule, medians, Sugeno integrals (SI), and ordinal weighted means (OWM), among others] and shows the difficulties for their application in this particular setting. Then, we propose two approaches to solve the drawbacks and we study their properties. Special emphasis is given to the study of monotonicity because the operator is proven nonsatisfactory for this property. Exhaustive empirical work shows that in most practical situations, this cannot be considered a problem. © 2003 Wiley Periodicals, Inc.
Josep Domingo-Ferrer, Vicenç Torra
Int. J. Intell. Syst.1
2003 Semantic-based aggregation for statistical disclosure control
abstract
In this paper we show how clustering can be used to aggregate different versions of the same data set in order to discover confidential information. Having these tools helps to not publish data that could be reidentified, which is known as Statistical Disclosure Control. In particular, the paper is focused on the case of dealing with categorical values. © 2003 Wiley Periodicals, Inc.
Aïda Valls, Vicenç Torra, Josep Domingo-Ferrer
Int. J. Intell. Syst.3
2003 On the connections between statistical disclosure control for microdata and some artificial intelligence tools
Josep Domingo-Ferrer, Vicenç Torra
Inf. Sci.1
2003 Collusion-secure and cost-effective detection of unlawful multimedia redistribution
abstract
Intellectual property protection of multimedia content is essential to the successful deployment of Internet content delivery platforms. There are two general approaches to multimedia copy protection: copy prevention and copy detection. Past experience shows that only copy detection based on mark embedding techniques looks promising. Multimedia fingerprinting means embedding a different buyer-identifying mark in each copy of the multimedia content being sold. Fingerprinting is subject to collusion attacks: a coalition of buyers collude and follow some strategy to mix their copies with the aim of obtaining a mixture from which none of their identifying marks can be retrieved; if their strategy is successful, the colluders can redistribute the mixture with impunity. A construction is presented in this paper to obtain fingerprinting codes for copyright protection which survive any collusion strategy involving up to three buyers (3-security). It is shown that the proposed scheme achieves 3-security with a codeword length dramatically shorter than the one required by the general Boneh-Shaw construction. Thus the proposed fingerprints require much less embedding capacity. Due to their own clandestine nature, collusions tend to involve a small number of buyers, so that there is plenty of use for codes providing cost-effective protection against collusions of size up to three.
Francesc Sebé, Josep Domingo-Ferrer
IEEE Trans. Syst. Man Cybern. Part C2
2002 Short 3-Secure Fingerprinting Codes for Copyright Protection
Francesc Sebé, Josep Domingo-Ferrer
ACISP2
2002 MICROCAST: Smart Card Based (Micro)Pay-per-View for Multicast Services
Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Francesc Sebé
CARDIS1
2002 A Provably Secure Additive and Multiplicative Privacy Homomorphism
Josep Domingo-Ferrer
ISC1
2002 Information-Theoretic Disclosure Risk Measures in Statistical Disclosure Control of Tabular Data
abstract
Statistical database protection is a part of information security which tries to prevent published statistical information (tables, individual records) from disclosing the contribution of specific respondents. This paper shows how to use information-theoretic concepts to measure disclosure risk for tabular data. The proposed disclosure risk measure is compatible with a broad class of disclosure protection methods and can be extended for computing disclosure risk for a set of linked tables.
Josep Domingo-Ferrer, Anna Oganian, Vicenç Torra
SSDBM1
2002 On the Security of Microaggregation with Individual Ranking: Analytical Attacks
abstract
Microaggregation is a statistical disclosure control technique. Raw microdata (i.e. individual records) are grouped into small aggregates prior to publication. With fixed-size groups, each aggregate contains k records to prevent disclosure of individual information. Individual ranking is a usual criterion to reduce multivariate microaggregation to univariate case: the idea is to perform microaggregation independently for each variable in the record. Using distributional assumptions, we show in this paper how to find interval estimates for the original data based on the microaggregated data. Such intervals can be considerably narrower than intervals resulting from subtraction of means, and can be useful to detect lack of security in a microaggregated data set. Analytical arguments given in this paper confirm recent empirical results about the unsafety of individual ranking microaggregation.
Josep Domingo-Ferrer, Josep Maria Mateo-Sanz, Anna Oganian, Aánge Torres
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2002 A Critique of the Sensitivity Rules Usually Employed for Statistical Table Protection
abstract
In statistical disclosure control of tabular data, sensitivity rules are commonly used to decide whether a table cell is sensitive and should therefore not be published. The most popular sensitivity rules are the dominance rule, the p%-rule and the pq-rule. The dominance rule has received critiques based on specific numerical examples and is being gradually abandoned by leading statistical agencies. In this paper, we construct general counterexamples which show that none of the above rules does adequately reflect disclosure risk if cell contributors or coalitions of them behave as intruders: in that case, releasing a cell declared non-sensitive can imply higher disclosure risk than releasing a cell declared sensitive. As possible solutions, we propose an alternative sensitivity rule based on the concentration of relative contributions. More generally, we suggest to complement a priori risk assessment based on sensitivity rules with a posteriori risk assessment which takes into account tables after they have been protected.
Josep Domingo-Ferrer, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2002 Editorial: Trends in Aggregation and Security Assessment for Inference Control in Statistical Databases
abstract
As e-commerce and Internet-based data handling become pervasive, companies and statistical agencies have the need to exploit the data they accumulate without violating citizens' privacy. Inference control is a discipline whose goal is to prevent published/exchanged data from being linked with the individual respondents they originated from. This special issue illustrates that inference control largely draws on soft computing and artificial intelligence techniques.
Vicenç Torra, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2002 Practical Data-Oriented Microaggregation for Statistical Disclosure Control
abstract
Microaggregation is a statistical disclosure control technique for microdata disseminated in statistical databases. Raw microdata (i.e., individual records or data vectors) are grouped into small aggregates prior to publication. Each aggregate should contain at least k data vectors to prevent disclosure of individual information, where k is a constant value preset by the data protector. No exact polynomial algorithms are known to date to microaggregate optimally, i.e., with minimal variability loss. Methods in the literature rank data and partition them into groups of fixed-size; in the multivariate case, ranking is performed by projecting data vectors onto a single axis. In this paper, candidate optimal solutions to the multivariate and univariate microaggregation problems are characterized. In the univariate case, two heuristics based on hierarchical clustering and genetic algorithms are introduced which are data-oriented in that they try to preserve natural data aggregates. In the multivariate case, fixed-size and hierarchical clustering microaggregation algorithms are presented which do not require data to be projected onto a single dimension; such methods clearly reduce variability loss as compared to conventional multivariate microaggregation on projected data.
Josep Domingo-Ferrer, Josep Maria Mateo-Sanz
IEEE Trans. Knowl. Data Eng.1
2001 Oblivious Image Watermarking Robust against Scaling and Geometric Distortions
Francesc Sebé, Josep Domingo-Ferrer
ISC2
2001 Current directions in smart cards
Josep Domingo-Ferrer, Pieter H. Hartel
Comput. Networks1
2000 A Performance Comparison of Java Cards for Micropayment Implementation
Jordi Castellà-Roca, Josep Domingo-Ferrer, Jordi Herrera-Joancomartí, Jordi Planes
CARDIS2
1998 Efficient Smart-Card Based Anonymous Fingerprinting
Josep Domingo-Ferrer, Jordi Herrera-Joancomartí
CARDIS1
1997 An implementable scheme for secure delegation of computing and data
Josep Domingo-Ferrer, Ricardo X. Sanchez del Castillo
ICICS1
1997 Multi-application smart cards and encrypted data, processing
Josep Domingo-Ferrer
Future Gener. Comput. Syst.1
1996 Multi-Application Smart Cards and Encrypted Data Processing
Josep Domingo-Ferrer
CARDIS1
1996 Achieving Rights Untransferability with Client-Independent Servers
Josep Domingo-Ferrer
Des. Codes Cryptogr.1
1996 A New Privacy Homomorphism and Applications
Josep Domingo-Ferrer
Inf. Process. Lett.1
1991 Algorithm- sequenced access control
Josep Domingo-Ferrer
Comput. Secur.1
1991 Distributed User Identification by Zero-Knowledge Access Rights Proving
Josep Domingo-Ferrer
Inf. Process. Lett.1
1990 Secure network bootstrapping: An algorithm for authentic key exchange and digital signitures
Josep Domingo-Ferrer, Llorenç Huguet i Rotger
Comput. Secur.1