David Sánchez 0001

dblp:45/445-1 · DBLP profile ↗
← Back
122ranked-venue papers
36as first author
29since 2021 · last 2026
0000-0001-7275-7887ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 58 · 15 first-author · 15 since 2021Databases, data management, data science and information retrieval · 32 · 13 first-author · 5 since 2021Security and privacy · 14 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 1 since 2021Computer networks · 8 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
YearPublicationVenuePosition
2026 FedVendorBC: Accurate Multi-vendor Federated Learning for Privacy-Preserving and Generalizable Breast Cancer Diagnosis Using Ultrasound Images
David Sánchez 0001, Zouhair Haddi, Josep Domingo-Ferrer
AIME (1)2
2026 Explainability-Driven Image Anonymization in Latent Space (EDIALS)
abstract
Facial image anonymization is essential to enable privacy-preserving image data sharing. The core challenge lies in removing identity-revealing information without degrading the utility of the images, which is essential for, e.g., demographic analysis. However, existing techniques apply uniform pixel-level distortions or synthesize replacements using Generative Adversarial Networks (GANs), which do not retain the meaningful features necessary for downstream tasks. To address this issue, we introduce EDIALS (Explainability-Driven Image Anonymization in Latent Space), which selectively modifies identity-specific latent features identified via explainability techniques. By applying targeted, incremental distortions in the latent space of an adversarial autoencoder, EDIALS effectively anonymizes images while preserving their analytical utility much better than existing techniques. Empirical evaluations on a common dataset show that EDIALS achieves 0.42% re-identification risk (equivalent to random guessing) while maintaining high utility: 84.66% F1 for age, 97.91% for gender, and 82.61% for race classification. In contrast, DeepPrivacy2 —a state-of-the-art GAN-based approach— results in a re-identification risk as large as 16.27% and lower utility: 78.32% F1, 82.58%, and 73.32% for age, gender, and race classification, respectively.
Younas Khan, Anna Monreale, Carlo Metta, David Sánchez 0001, Josep Domingo-Ferrer
CODASPY4
2026 A comparative analysis, enhancement and evaluation of text anonymization with pre-trained Large Language Models
abstract
Large Language Models (LLMs) have gained prominence for their remarkable proficiency across various natural language processing tasks. Recent studies have suggested their potential to outperform current text anonymization methods, although an objective evaluation is needed to validate these claims. To address this issue, this work introduces a comprehensive evaluation framework that automatically assesses both privacy protection and utility preservation without relying on manually curated ground-truth data. Moreover, we conduct an in-depth analysis of the LLM-based text anonymization methods proposed so far. Building on the strengths and limitations we found, we propose a novel method to enhance anonymization quality. We also report extensive experimental comparisons between LLM-based approaches and a variety of previous techniques, including those based on named entity recognition (NER), and those more oriented towards privacy-preserving data publishing (PPDP). The results show that LLM-based approaches effectively outperform traditional methods in terms of privacy and utility. Furthermore, we benchmark against manual anonymization, which performed poorly, thus highlighting the limitations of using them as evaluation ground truth. Notably, our LLM-based method stood out by achieving the best privacy protection, and the best privacy-utility trade-off.
Benet Manzanares-Salor, David Sánchez 0001
Expert Syst. Appl.2
2026 Unsupervised utility evaluation of text anonymization methods via neural language models
abstract
Text anonymization methods strive to find a balance between privacy protection and utility preservation, where the latter refers to the fact that the anonymized documents are still analytically useful for research or business tasks. The performance of these methods is evaluated empirically, by comparing their outputs with human-based anonymizations through the standard precision and recall metrics. Whereas recall is used as a proxy for the level of privacy attained, precision is interpreted as the amount of utility preserved. Nonetheless, these metrics were not designed for evaluating privacy-oriented tasks and present several drawbacks. First, they assume a unique ground truth whereas, in text anonymization, several masking choices can be equally valid to prevent re-identification. Second, the human annotations used as ground truth are inherently subjective and prone to errors. Third, both metrics weight terms uniformly, thereby ignoring the varied impact that the terms' semantics have on utility and re-identification risk. To overcome these limitations, in this paper we present the first unsupervised utility metric for anonymized texts. Our metric relies on neural language models to quantify the utility loss incurred by anonymization. Empirical experiments on document clustering show that our metric captures the actual utility of the anonymized outcomes more accurately than precision, while being more sensitive to varying anonymization intensities. Together with a previously proposed privacy metric, our proposal defines a complete evaluation framework for text anonymization that does not require costly human annotations. Using this framework, we also report a comprehensive evaluation of a variety of text anonymization methods.
Benet Manzanares-Salor, David Sánchez 0001, Pierre Lison
Neural Networks2
2026 Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions
abstract
Membership inference attacks (MIAs) have become the standard tool for evaluating privacy leakage in machine learning (ML). Among them, the Likelihood-Ratio Attack (LiRA) is widely regarded as the state of the art when sufficient shadow models are available. However, prior evaluations have often overstated the effectiveness of LiRA by attacking models overconfident on their training samples, calibrating thresholds on target data, assuming balanced membership priors, and/or overlooking attack reproducibility. We re-evaluate LiRA under a realistic protocol that (i) trains models using anti-overfitting (AOF) (and transfer learning (TL), when applicable) to reduce overconfidence as it would be desirable in production models; (ii) calibrates decision thresholds from shadow models and data rather than (usually unavailable) target data; (iii) measures positive predictive value (PPV, a.k.a. precision) under shadow-based thresholds and skewed - rather than unrealistically balanced - membership priors (pi <= 10%); and (iv) quantifies per-sample membership reproducibility across different seeds and training variations. In this setting, we find that (a) AOF significantly weakens LiRA and TL further reduces the effectiveness of the attack, while improving model accuracy; (b) with shadow-based thresholds and skewed priors, LiRA's PPV often drops from near-perfect to substantially lower levels, especially under AOF/AOF+TL and for pi <= 10%; and (c) LiRA's thresholded vulnerable sets at extremely low FPR exhibit poor reproducibility across runs, while likelihood ratio-based rankings are more stable. These results suggest that (i) LiRA, and likely weaker MIAs, are less effective than previously suggested, and their positive inferences can be less reliable under realistic settings; and (ii) for MIAs to serve as meaningful privacy auditing tools, their evaluation must reflect pragmatic training practices, feasible attacker assumptions, and reproducibility considerations. We release our code at: https://github.com/najeebjebreel/lira_analysis.
Najeeb Jebreel, Mona Khalil, David Sánchez 0001, Josep Domingo-Ferrer
Proc. Priv. Enhancing Technol.3
2026 IKXAI-anonymity: Iterative XAI-based probabilistic k-anonymity for face image anonymization
abstract
Facial images are of great interest for research, but they might compromise individuals’ privacy. Existing approaches to image anonymization often rely on either uniform image perturbation (such as pixelation or blurring) or generative AI models, but both lack privacy guarantees and may severely compromise images’ analytical utility. In this work, we present IKXAI-anonymity, a novel framework that enforces probabilistic k -anonymity on face images through an iterative process guided by explainable artificial intelligence techniques. Unlike global transformations or latent-space manipulations, our method applies small and incremental modifications to only the most identity-revealing pixels iteratively. The added noise remains lightweight at every step, which allows the anonymization process to gradually reduce identity traces with minimal harm the overall utility of the image in secondary analyses ( e.g. , age, gender, and race classification). The identity-revealing image regions to perturb are detected using a gradient-based counterfactual explainer. The process iterates until the image identity prediction rank falls below a threshold k , which corresponds to a probability of re-identification of at most 1 / k . We evaluate IKXAI-anonymity on a curated dataset of face images and compare it with pixel perturbation techniques, direct image-space k -anonymity and state-of-the-art generative approaches. Our results show that IKXAI-anonymity reliably achieves the targeted level of privacy while retaining much better image utility than the other methods in secondary tasks. The code to reproduce all the reported experiments is available at https://github.com/RamiHaf/IKXAI-anonymity .
Rami Haffar, David Sánchez 0001
Pattern Recognit.2
2025 Privacy- & Utility-Preserving Data Releases over Fragmented Data Using Individual Differential Privacy
Luis Del Vasto-Terrientes, Sergio Martínez, David Sánchez 0001
ICISSP (2)3
2025 Defenses Against Membership Inference Attacks on Unlearned Data
Josep Domingo-Ferrer, Najeeb Jebreel, David Sánchez 0001
MDAI3
2025 Enhancing efficiency and data utility in longitudinal data anonymization
Fatemeh Amiri, David Sánchez 0001, Josep Domingo-Ferrer
Inf. Sci.2
2025 Enhancing text anonymization via re-identification risk-based explainability
Benet Manzanares-Salor, David Sánchez 0001
Knowl. Based Syst.2
2025 MemberShield: A framework for federated learning with membership privacy
abstract
Federated Learning (FL) allows multiple data owners to build high-quality deep learning models collaboratively, by sharing only model updates and keeping data on their premises. Even though FL offers privacy-by-design, it is vulnerable to membership inference attacks (MIA), where an adversary tries to determine whether a sample was included in the training data. Existing defenses against MIA cannot offer meaningful privacy protection without significantly hampering the model's utility and causing a non-negligible training overhead. In this paper we analyze the underlying causes of the differences in the model behavior for member and non-member samples, which arise from model overfitting and facilitate MIAs. Accordingly, we propose MemberShield, a generalization-based defense method for MIAs that consists of: (i) one-time preprocessing of each client's training data labels that transforms one-hot encoded labels to soft labels and eventually exploits them in local training, and (ii) early stopping the training when the local model's validation accuracy does not improve on that of the global model for a number of epochs. Extensive empirical evaluations on three widely used datasets and four model architectures demonstrate that MemberShield outperforms state-of-the-art defense methods by delivering substantially better practical privacy protection against all forms of MIAs, while better preserving the target model utility. On top of that, our proposal significantly reduces training time and is straightforward to implement, by just tuning a single hyperparameter.
David Sánchez 0001, Zouhair Haddi, Josep Domingo-Ferrer
Neural Networks2
2025 DP2Unlearning: An efficient and guaranteed unlearning framework for LLMs
abstract
Large language models (LLMs) have recently revolutionized language processing tasks but have also brought ethical and legal issues. LLMs have a tendency to memorize potentially private or copyrighted information present in the training data, which might then be delivered to end users at inference time. When this happens, a naive solution is to retrain the model from scratch after excluding the undesired data. Although this guarantees that the target data have been forgotten, it is also prohibitively expensive for LLMs. Approximate unlearning offers a more efficient alternative, as it consists of ex post modifications of the trained model itself to prevent undesirable results, but it lacks forgetting guarantees because it relies solely on empirical evidence. In this work, we present DP2Unlearning, a novel LLM unlearning framework that offers formal forgetting guarantees at a significantly lower cost than retraining from scratch on the data to be retained. DP2Unlearning involves training LLMs on textual data protected using ϵ-differential privacy (DP), which later enables efficient unlearning with the guarantees against disclosure associated with the chosen ϵ. Our experiments demonstrate that DP2Unlearning achieves similar model performance post-unlearning, compared to an LLM retraining from scratch on retained data -the gold standard exact unlearning- but at approximately half the unlearning cost. In addition, with a reasonable computational cost, it outperforms approximate unlearning methods at both preserving the utility of the model post-unlearning and effectively forgetting the targeted information. The code of our experiments is available at https://github.com/tamimalmahmud/LLM-Unlearning/tree/main/DP2Unlearning.
Tamim Al Mahmud, Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001
Neural Networks4
2024 An Examination of the Alleged Privacy Threats of Confidence-Ranked Reconstruction of Census Microdata
David Sánchez 0001, Najeeb Jebreel, Krishnamurty Muralidhar, Josep Domingo-Ferrer, Alberto Blanco-Justicia
PSD1
2024 Evaluating the disclosure risk of anonymized documents via a machine learning-based re-identification attack
abstract
Abstract The availability of textual data depicting human-centered features and behaviors is crucial for many data mining and machine learning tasks. However, data containing personal information should be anonymized prior making them available for secondary use. A variety of text anonymization methods have been proposed in the last years, which are standardly evaluated by comparing their outputs with human-based anonymizations. The residual disclosure risk is estimated with the recall metric, which quantifies the proportion of manually annotated re-identifying terms successfully detected by the anonymization algorithm. Nevertheless, recall is not a risk metric, which leads to several drawbacks. First, it requires a unique ground truth, and this does not hold for text anonymization, where several masking choices could be equally valid to prevent re-identification. Second, it relies on human judgements, which are inherently subjective and prone to errors. Finally, the recall metric weights terms uniformly, thereby ignoring the fact that the influence on the disclosure risk of some missed terms may be much larger than of others. To overcome these drawbacks, in this paper we propose a novel method to evaluate the disclosure risk of anonymized texts by means of an automated re-identification attack. We formalize the attack as a multi-class classification task and leverage state-of-the-art neural language models to aggregate the data sources that attackers may use to build the classifier. We illustrate the effectiveness of our method by assessing the disclosure risk of several methods for text anonymization under different attack configurations. Empirical results show substantial privacy risks for most existing anonymization methods.
Benet Manzanares-Salor, David Sánchez 0001, Pierre Lison
Data Min. Knowl. Discov.2
2024 LFighter: Defending against the label-flipping attack in federated learning
Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001, Alberto Blanco-Justicia
Neural Networks3
2024 Enhanced Security and Privacy via Fragmented Federated Learning
abstract
In federated learning (FL), a set of participants share updates computed on their local data with an aggregator server that combines updates into a global model. However, reconciling accuracy with privacy and security is a challenge to FL. On the one hand, good updates sent by honest participants may reveal their private local information, whereas poisoned updates sent by malicious participants may compromise the model's availability and/or integrity. On the other hand, enhancing privacy via update distortion damages accuracy, whereas doing so via update aggregation damages security because it does not allow the server to filter out individual poisoned updates. To tackle the accuracy-privacy-security conflict, we propose fragmented FL (FFL), in which participants randomly exchange and mix fragments of their updates before sending them to the server. To achieve privacy, we design a lightweight protocol that allows participants to privately exchange and mix encrypted fragments of their updates so that the server can neither obtain individual updates nor link them to their originators. To achieve security, we design a reputation-based defense tailored for FFL that builds trust in participants and their mixed updates based on the quality of the fragments they exchange and the mixed updates they send. Since the exchanged fragments' parameters keep their original coordinates and attackers can be neutralized, the server can correctly reconstruct a global model from the received mixed updates without accuracy loss. Experiments on four real data sets show that FFL can prevent semi-honest servers from mounting privacy attacks, can effectively counter-poisoning attacks, and can keep the accuracy of the global model.
Najeeb Jebreel, Josep Domingo-Ferrer, Alberto Blanco-Justicia, David Sánchez 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Explaining predictions and attacks in federated learning via random forests
abstract
Abstract Artificial intelligence (AI) is used for various purposes that are critical to human life. However, most state-of-the-art AI algorithms are black-box models, which means that humans cannot understand how such models make decisions. To forestall an algorithm-based authoritarian society, decisions based on machine learning ought to inspire trust by being explainable . For AI explainability to be practical, it must be feasible to obtain explanations systematically and automatically. A usual methodology to explain predictions made by a (black-box) deep learning model is to build a surrogate model based on a less difficult, more understandable decision algorithm. In this work, we focus on explaining by means of model surrogates the (mis)behavior of black-box models trained via federated learning. Federated learning is a decentralized machine learning technique that aggregates partial models trained by a set of peers on their own private data to obtain a global model. Due to its decentralized nature, federated learning offers some privacy protection to the participating peers. Nonetheless, it remains vulnerable to a variety of security attacks and even to sophisticated privacy attacks. To mitigate the effects of such attacks, we turn to the causes underlying misclassification by the federated model, which may indicate manipulations of the model. Our approach is to use random forests containing decision trees of restricted depth as surrogates of the federated black-box model. Then, we leverage decision trees in the forest to compute the importance of the features involved in the wrong predictions. We have applied our method to detect security and privacy attacks that malicious peers or the model manager may orchestrate in federated learning scenarios. Empirical results show that our method can detect attacks with high accuracy and, unlike other attack detection mechanisms, it can also explain the operation of such attacks at the peers’ side.
Rami Haffar, David Sánchez 0001, Josep Domingo-Ferrer
Appl. Intell.2
2023 Secure, accurate and privacy-aware fully decentralized learning via co-utility
abstract
Fully decentralized learning is a setting in which each peer in a P2P network trains a machine learning model with the help of the other peers. Each peer acts as a model manager by periodically sending her current model to other peers, who answer by returning model updates they compute on their private data. This creates a tension among privacy, accuracy and security. The privacy risk is that model updates returned by a peer can leak some of the peer’s private data. Unfortunately, distorting model updates to protect privacy works against the accuracy of the trained model. On the other hand, aggregating the updates of several peers and then sending the aggregate to the model manager may preserve privacy but it goes against security, because the model manager cannot filter out individual bad updates. Also, peers are autonomous and hence it cannot be taken for granted that they will honestly supply model updates to help the model manager train her model. To reconcile accuracy, privacy and security, we present a fully decentralized learning protocol such that: (i) it allows perfectly accurate individual updates to be returned by peers to the model manager in a privacy-preserving manner; (ii) it is co-utile by design, that is, it incentivizes rational peers to follow the protocol without deviating. The latter feature discourages rational attacks that might compromise security and it also deters free riding, thereby ensuring the sustainability of the protocol.
Jesús A. Manjón, Josep Domingo-Ferrer, David Sánchez 0001, Alberto Blanco-Justicia
Comput. Commun.3
2023 Utility-Preserving Privacy Protection of Textual Documents via Word Embeddings
abstract
A great variety of mechanisms have been proposed to protect structured databases with numerical and categorical attributes; however, little attention has been devoted to unstructured textual data. Textual data protection requires first detecting sensitive pieces of text and then masking those pieces via suppression or generalization. Current solutions rely on classifiers that can recognize a fixed set of (allegedly sensitive) named entities. Yet, such approaches fall short of providing adequate protection because in reality references to sensitive information are not limited to a predefined set of entity types, and not all the appearances of certain entity type result in disclosure. In this work we propose a more general and flexible based on the notion of word embedding. By means of word embeddings we build vectors that numerically capture the semantic relationships of the textual terms. Then we evaluate the disclosure caused by the terms on the entity to be protected according to the similarity between their vector representations. Our method also preserves the semantics (and, therefore, the utility) of the document by replacing risky terms with privacy-preserving generalizations. Empirical results show that our approach offers much more robust protection and greater utility preservation than methods based on named entity recognition.
Fadi Hassan, David Sánchez 0001, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.2
2022 Automatic Evaluation of Disclosure Risks of Text Anonymization Methods
Benet Manzanares-Salor, David Sánchez 0001, Pierre Lison
PSD2
2022 The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
abstract
Abstract We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods. Text anonymization, defined as the task of editing a text document to prevent the disclosure of personal information, currently suffers from a shortage of privacy-oriented annotated text resources, making it difficult to properly evaluate the level of privacy protection offered by various anonymization methods. This paper presents TAB (Text Anonymization Benchmark), a new, open-source annotated corpus developed to address this shortage. The corpus comprises 1,268 English-language court cases from the European Court of Human Rights (ECHR) enriched with comprehensive annotations about the personal information appearing in each document, including their semantic category, identifier type, confidential attributes, and co-reference relations. Compared with previous work, the TAB corpus is designed to go beyond traditional de-identification (which is limited to the detection of predefined semantic categories), and explicitly marks which text spans ought to be masked in order to conceal the identity of the person to be protected. Along with presenting the corpus and its annotation layers, we also propose a set of evaluation metrics that are specifically tailored toward measuring the performance of text anonymization, both in terms of privacy protection and utility preservation. We illustrate the use of the benchmark and the proposed metrics by assessing the empirical performance of several baseline text anonymization models. The full corpus along with its privacy-oriented annotation guidelines, evaluation scripts, and baseline models are available on: https://github.com/NorskRegnesentral/text-anonymization-benchmark.
Ildikó Pilán, Pierre Lison, Lilja Øvrelid, Anthi Papadopoulou, David Sánchez 0001, Montserrat Batet
Comput. Linguistics5
2022 Decentralized k-anonymization of trajectories via privacy-preserving tit-for-tat
abstract
Mobility data, and specifically trajectories, are used to monitor the mobility of the population and are crucial to improve public health, transportation, urban planning, economic planning, etc. However, trajectories are personally identifiable information and hence they should be anonymized before releasing them for secondary use. Anonymization cannot be limited to suppressing the metadata containing the subject’s identity, because the origin, the destination and even the intermediate points of a trajectory may allow re-identifying the subject who followed it. Proper anonymization requires masking detailed spatiotemporal information. The standard approach to build anonymized data sets is centralized: the subjects send their original movement data to a controller, who takes care of producing an anonymized mobility data set. This requires subjects to blindly trust the controller. In this paper, we empower subjects with the ability to anonymize their trajectories locally by adhering to a privacy model in order to achieve formal privacy guarantees. After reviewing the state of the art, we motivate our choice of k-anonymity as a privacy model. We then set out to decentralize k-anonymity in a rational setting: a subject k-anonymizes her completed trajectory by aggregating with k−1 similar trajectories obtained from other (unknown) subjects. The latter trajectories are gathered via an anonymous and privacy-preserving tit-for-tat data exchange protocol, which runs on a fully decentralized peer-to-peer network. Experiments show that, without relying on a (trusted) data controller and while ensuring privacy w.r.t. other peers, our approach yields k-anonymized mobility data sets that are still reasonably useful compared to the near-optimal data sets obtained in the centralized approach.
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001
Comput. Commun.3
2022 Generating Deep Learning Model-Specific Explanations at the End User's Side
abstract
End users who cannot afford to collect and label big data to train accurate deep learning (DL) models resort to Machine Learning as a Service (MLaaS) providers, who provide paid access to accurate DL models. However, the lack of transparency in how the providers’ models make predictions causes a problem of trust. A way to increase trust (and also to align with ethical regulations) is for predictions to be accompanied by explanations locally and independently generated by the end users (rather than by explanations offered by the model providers). Explanation methods using internal components of DL models (a.k.a. model-specific explanations) are more accurate and effective than those relying solely on the inputs and outputs (a.k.a. model-agnostic explanations). However, end users lack white-box access to the internal components of the providers’ models. To tackle this issue, we propose a novel approach allowing an end user to locally generate model-specific explanations for a DL classification model accessed via a provider’s API. First, we approximate the provider’s model with a local surrogate model. We then use the surrogate model’s components to locally generate model-specific explanations that approximate the explanations obtainable with white-box access to the provider’s DL model. Specifically, we leverage the surrogate model’s gradients to generate adversarial examples that counterfactually explain why an input example is classified into a specific class. Our approach only requires the end user to have unlabeled data of size [Formula: see text] of the provider’s training data and with a similar distribution; given the small size and unlabeled nature of these data, they can be assumed to be already available to the end user or even to be supplied by the provider to build trust in his model. We demonstrate the accuracy and effectiveness of our approach through extensive experiments on two ML tasks: image classification and tabular data classification. The locally generated explanations are consistent with those obtainable with white-box access to the provider’s model, thus giving end users an independent and reliable way to determine if the provider’s model is trustworthy.
Rami Haffar, Najeeb Jebreel, David Sánchez 0001, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2022 Secure and Privacy-Preserving Federated Learning via Co-Utility
abstract
The decentralized nature of federated learning, that often leverages the power of edge devices, makes it vulnerable to attacks against privacy and security. The privacy risk for a peer is that the model update she computes on her private data may, when sent to the model manager, leak information on those private data. Even more obvious are security attacks, whereby one or several malicious peers return wrong model updates in order to disrupt the learning process and lead to a wrong model being learned. In this article, we build a federated learning framework that offers privacy to the participating peers as well as security against the Byzantine and poisoning attacks. Our framework consists of several protocols that provide strong privacy to the participating peers via unlinkable anonymity and that are rationally sustainable based on the co-utility property. In other words, no rational party is interested in deviating from the proposed protocols. We leverage the notion of co-utility to build a decentralized co-utile reputation management system that provides incentives for parties to adhere to the protocols. Unlike privacy protection via differential privacy, our approach preserves the values of model updates and, hence, the accuracy of plain federated learning; unlike privacy protection via update aggregation, our approach preserves the ability to detect bad model updates while substantially reducing the computational overhead compared to methods based on homomorphic encryption.
Josep Domingo-Ferrer, Alberto Blanco-Justicia, Jesús A. Manjón, David Sánchez 0001
IEEE Internet Things J.4
2021 Anonymisation Models for Text Data: State of the art, Challenges and Future Directions
abstract
Pierre Lison, Ildikó Pilán, David Sanchez, Montserrat Batet, Lilja Øvrelid. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Pierre Lison, Ildikó Pilán, David Sánchez 0001, Montserrat Batet, Lilja Øvrelid
ACL/IJCNLP (1)3
2021 Explaining Image Misclassification in Deep Learning via Adversarial Examples
Rami Haffar, Najeeb Jebreel, Josep Domingo-Ferrer, David Sánchez 0001
MDAI4
2021 Achieving security and privacy in federated learning systems: Survey, research challenges and future directions
abstract
Federated learning (FL) allows a server to learn a machine learning (ML) model across multiple decentralized clients that privately store their own training data. In contrast with centralized ML approaches, FL saves computation to the server and does not require the clients to outsource their private data to the server. However, FL is not free of issues. On the one hand, the model updates sent by the clients at each training epoch might leak information on the clients’ private data. On the other hand, the model learnt by the server may be subjected to attacks by malicious clients; these security attacks might poison the model or prevent it from converging. In this paper, we first examine security and privacy attacks to FL and critically survey solutions proposed in the literature to mitigate each attack. Afterwards, we discuss the difficulty of simultaneously achieving security and privacy protection. Finally, we sketch ways to tackle this open problem and attain both security and privacy.
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Adrian Flanagan, Kuan Eeik Tan
Eng. Appl. Artif. Intell.4
2021 A large reproducible benchmark of ontology-based methods and word embeddings for word similarity
Juan J. Lastra-Díaz, Josu Goikoetxea, Mohamed Ali Hadj Taieb, Ana García-Serrano, Mohamed Benaouicha, Eneko Agirre, David Sánchez 0001
Inf. Syst.7
2021 Privacy protection of user profiles in online search via semantic randomization
Mercedes Rodriguez-Garcia, Montserrat Batet, David Sánchez 0001, Alexandre Viejo
Knowl. Inf. Syst.3
2020 Co-Utile Peer-to-Peer Decentralized Computing
abstract
Outsourcing computation allows wielding huge computational power. Even though cloud computing is the most usual type of outsourcing, resorting to idle edge devices for decentralized computation is an increasingly attractive alternative. We tackle the problem of making peer honesty and thus computation correctness self-enforcing in decentralized computing with untrusted peers. To do so, we leverage the co-utility property, which characterizes a situation in which honest co-operation is the best rational option to take even for purely selfish agents; in particular, if a protocol is co-utile, it is self-enforcing. Reputation is a powerful incentive that can make a P2P protocol co-utile. We present a co-utile P2P decentralized computing protocol that builds on a decentralized reputation calculation, which is itself co-utile and therefore self-enforcing. In this protocol, peers are given a computational task including code and data and they are incentivized to compute it correctly. Based also on co-utile reputation, we then present a protocol for federated learning, whereby peers compute on their local private data and have no incentive to randomly attack or poison the model. Our experiments show the viability of our co-utile approach to obtain correct results in both decentralized computation and federated learning.
Josep Domingo-Ferrer, Alberto Blanco-Justicia, David Sánchez 0001, Najeeb Jebreel
CCGRID3
2020 Fair Detection of Poisoning Attacks in Federated Learning
abstract
Federated learning is a decentralized machine learning technique that aggregates partial models trained by a set of clients on their own private data to obtain a global model. This technique is vulnerable to security attacks, such as model poisoning, whereby malicious clients submit bad updates in order to prevent the model from converging or to introduce artificial bias in the classification. Applying anti-poisoning techniques might lead to the discrimination of minority groups whose data are significantly and legitimately different from those of the majority of clients. In this work, we strive to strike a balance between fighting poisoning and accommodating diversity to help learning fairer and less discriminatory federated learning models. In this way, we forestall the exclusion of diverse clients while still ensuring detection of poisoning attacks. Empirical work on a standard machine learning data set shows that employing our approach to tell legitimate from malicious updates produces models that are more accurate than those obtained with standard poisoning detection techniques.
Ashneet Khandpur Singh, Alberto Blanco-Justicia, Josep Domingo-Ferrer, David Sánchez 0001, David Rebollo-Monedero
ICTAI4
2020 Explaining Misclassification and Attacks in Deep Learning via Random Forests
Rami Haffar, Josep Domingo-Ferrer, David Sánchez 0001
MDAI3
2020 Efficient Detection of Byzantine Attacks in Federated Learning Using Last Layer Biases
Najeeb Jebreel, Alberto Blanco-Justicia, David Sánchez 0001, Josep Domingo-Ferrer
MDAI3
2020 Detecting Bad Answers in Survey Data Through Unsupervised Machine Learning
Najeeb Jebreel, Rami Haffar, Ashneet Khandpur Singh, David Sánchez 0001, Josep Domingo-Ferrer, Alberto Blanco-Justicia
PSD4
2020 µ-ANT: semantic microaggregation-based anonymization tool
abstract
MOTIVATION: Detailed patient data are crucial for medical research. Yet, these healthcare data can only be released for secondary use if they have undergone anonymization. RESULTS: We present and describe µ-ANT, a practical and easily configurable anonymization tool for (healthcare) data. It implements several state-of-the-art methods to offer robust privacy guarantees and preserve the utility of the anonymized data as much as possible. µ-ANT also supports the heterogenous attribute types commonly found in electronic healthcare records and targets both practitioners and software developers interested in data anonymization. AVAILABILITY AND IMPLEMENTATION: (source code, documentation, executable, sample datasets and use case examples) https://github.com/CrisesUrv/microaggregation-based_anonymization_tool.
David Sánchez 0001, Sergio Martínez, Josep Domingo-Ferrer, Jordi Soria-Comas, Montserrat Batet
Bioinform.1
2020 Secure monitoring in IoT-based services via fog orchestration
Alexandre Viejo, David Sánchez 0001
Future Gener. Comput. Syst.2
2020 Outsourcing analyses on privacy-protected multivariate categorical data stored in untrusted clouds
Josep Domingo-Ferrer, David Sánchez 0001, Sara Ricci, Mónica Muñoz-Batista
Knowl. Inf. Syst.2
2020 Machine learning explainability via microaggregation and shallow decision trees
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001
Knowl. Based Syst.4
2019 Secure and privacy-preserving orchestration and delivery of fog-enabled IoT services
Alexandre Viejo, David Sánchez 0001
Ad Hoc Networks2
2019 Privacy-preserving cloud computing on sensitive data: A survey of methods, products and challenges
Josep Domingo-Ferrer, Oriol Farràs, Jordi Ribes-González, David Sánchez 0001
Comput. Commun.4
2018 Privacy-preserving and advertising-friendly web surfing
David Sánchez 0001, Alexandre Viejo
Comput. Commun.1
2018 Survey and evaluation of web search engine hit counts as research tools in computational linguistics
David Sánchez 0001, Laura Martínez-Sanahuja, Montserrat Batet
Inf. Syst.1
2018 A semantic-preserving differentially private method for releasing query logs
David Sánchez 0001, Montserrat Batet, Alexandre Viejo, Mercedes Rodriguez-Garcia, Jordi Castellà-Roca
Inf. Sci.1
2018 Co-utile disclosure of private data in social networks
David Sánchez 0001, Josep Domingo-Ferrer, Sergio Martínez
Inf. Sci.1
2017 Privacy-preserving data outsourcing in the cloud via semantic data splitting
David Sánchez 0001, Montserrat Batet
Comput. Commun.1
2017 Toward sensitive document release with privacy guarantees
David Sánchez 0001, Montserrat Batet
Eng. Appl. Artif. Intell.1
2017 Co-Utility: Self-Enforcing protocols for the mutual benefit of participants
Josep Domingo-Ferrer, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Eng. Appl. Artif. Intell.3
2017 A semantic framework for noise addition with nominal data
Mercedes Rodriguez-Garcia, Montserrat Batet, David Sánchez 0001
Knowl. Based Syst.3
2017 Individual Differential Privacy: A Utility-Preserving Formulation of Differential Privacy Guarantees
abstract
Differential privacy is a popular privacy model within the research community because of the strong privacy guarantee it offers, namely that the presence or absence of any individual in a data set does not significantly influence the results of analyses on the data set. However, enforcing this strict guarantee in practice significantly distorts data and/or limits data uses, thus diminishing the analytical utility of the differentially private results. In an attempt to address this shortcoming, several relaxations of differential privacy have been proposed that trade off privacy guarantees for improved data utility. In this paper, we argue that the standard formalization of differential privacy is stricter than required by the intuitive privacy guarantee it seeks. In particular, the standard formalization requires indistinguishability of results between any pair of neighbor data sets, while indistinguishability between the actual data set and its neighbor data sets should be enough. This limits the data controller's ability to adjust the level of protection to the actual data, hence resulting in significant accuracy loss. In this respect, we propose individual differential privacy, an alternative differential privacy notion that offers the same privacy guarantees as standard differential privacy to individuals (even though not to groups of individuals). This new notion allows the data controller to adjust the distortion to the actual data set, which results in less distortion and more analytical accuracy. We propose several mechanisms to attain individual differential privacy and we compare the new notion against standard differential privacy in terms of the accuracy of the analytical results.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, David Megías 0001
IEEE Trans. Inf. Forensics Secur.3
2016 Ontology-based Access Control Management: Two Use Cases
abstract
Access control management is an important area of research within the security field. Several models have been proposed to manage the access rights of users over restricted resources, which are mainly based on defining rules between specific entities and concrete resources. Though these approaches are enough to manage organizations involving a limited number of entities and resources, the specification of rules or constraints for large and heterogeneous scenarios may imply a considerable burden to the administrators. To palliate this problem, we propose a generic ontology-based solution to manage the access control that can greatly simplify and speed up the definition of rules in complex scenarios and that can also improve the interoperability between heterogeneous settings. Moreover, we show its potential by applying it in two highly dynamic and large scenarios, i.e., Online Social Networks (OSNs) and the Cloud.
Malik Imran Daud, David Sánchez 0001, Alexandre Viejo
ICAART (1)2
2016 t-closeness through microaggregation: Strict privacy with enhanced utility preservation
abstract
This paper proposes and shows how to use microaggregation to attain t-closeness on top of k-anonymity to protect data releases. The advantages in terms of data utility preservation of microaggregation over classic approaches based on generalizing values are analyzed. Then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
ICDE3
2016 Improving Semantic Relatedness Assessments: Ontologies Meet Textual Corpora
abstract
Even though the calculation of the semantic similarity between textual entities has received a lot of attention by the research community, the more general notion of semantic relatedness (which considers both taxonomic and non-taxonomic knowledge) has been significantly less studied and, in general, stays one step behind in terms of accuracy. In this paper, we improve semantic relatedness assessments by aggregating the highly-accurate ontology-based estimation of semantic similarity with the distributional resemblance of textual terms computed from large textual corpora. As a result, our approach is able to improve the accuracy of related works on a standard benchmark.
Montserrat Batet, David Sánchez 0001
KES2
2016 Evaluating the Suitability of Web Search Engines as Proxies for Knowledge Discovery from the Web
abstract
Many researchers use the Web search engines’ hit count as an estimator of the Web information distribution in a variety of knowledge-based (linguistic) tasks. Even though many studies have been conducted on the retrieval effectiveness of Web search engines for Web users, few of them have evaluated them as research tools. In this study we analyse the currently available search engines and evaluate the suitability and accuracy of the hit counts they provide as estimators of the frequency/probability of textual entities. From the results of this study, we identify the search engines best suited to be used in linguistic research.
Laura Martínez-Sanahuja, David Sánchez 0001
KES2
2016 Privacy-Preserving Cloud-Based Statistical Analyses on Sensitive Categorical Data
Sara Ricci, Josep Domingo-Ferrer, David Sánchez 0001
MDAI3
2016 Perturbative Data Protection of Multivariate Nominal Datasets
Mercedes Rodriguez-Garcia, David Sánchez 0001, Montserrat Batet
PSD2
2016 Privacy-driven access control in social networks by means of automatic semantic annotation
Malik Imran Daud, David Sánchez 0001, Alexandre Viejo
Comput. Commun.2
2016 Enforcing transparent access to private content in social networks by means of automatic sanitization
abstract
Social networks have become an essential meeting point for millions of individuals willing to publish and consume huge quantities of heterogeneous information. Some studies have shown that the data published in these platforms may contain sensitive personal information and that external entities can gather and exploit this knowledge for their own benefit. Even though some methods to preserve the privacy of social networks users have been proposed, they generally apply rigid access control measures to the protected content and, even worse, they do not enable the users to understand which contents are sensitive. Last but not least, most of them require the collaboration of social network operators or they fail to provide a practical solution capable of working with well-known and already deployed social platforms. In this paper, we propose a new scheme that addresses all these issues. The new system is envisaged as an independent piece of software that does not depend on the social network in use and that can be transparently applied to most existing ones. According to a set of privacy requirements intuitively defined by the users of a social network, the proposed scheme is able to: (i) automatically detect sensitive data in users’ publications; (ii) construct sanitized versions of such data; and (iii) provide privacy-preserving transparent access to sensitive contents by disclosing more or less information to readers according to their credentials toward the owner of the publications. We also study the applicability of the proposed system in general and illustrate its behavior in two case studies.
Alexandre Viejo, David Sánchez 0001
Expert Syst. Appl.2
2016 Self-enforcing protocols via co-utile reputation management
Josep Domingo-Ferrer, Oriol Farràs, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Inf. Sci.4
2016 C-sanitized: A privacy model for document redaction and sanitization
abstract
Vast amounts of information are daily exchanged and/or released. The sensitive nature of much of this information creates a serious privacy threat when documents are uncontrollably made available to untrusted third parties. In such cases, appropriate data protection measures should be undertaken by the responsible organization, especially under the umbrella of current legislation on data privacy. To do so, human experts are usually requested to redact or sanitize document contents. To relieve this burdensome task, this paper presents a privacy model for document redaction/sanitization, which offers several advantages over other models available in the literature. Based on the well‐established foundations of data semantics and information theory, our model provides a framework to develop and implement automated and inherently semantic redaction/sanitization tools. Moreover, contrary to ad‐hoc redaction methods, our proposal provides a priori privacy guarantees which can be intuitively defined according to current legislations on data privacy. Empirical tests performed within the context of several use cases illustrate the applicability of our model and its ability to mimic the reasoning of human sanitizers.
David Sánchez 0001, Montserrat Batet
J. Assoc. Inf. Sci. Technol.1
2015 Ontology Selection for Semantic Similarity Assessment
Montserrat Batet, David Sánchez 0001
ICAART (2)2
2015 Privacy Risk Assessment of Textual Publications in Social Networks
David Sánchez 0001, Alexandre Viejo
ICAART (1)1
2015 Semantic Noise: Privacy-Protection of Nominal Microdata through Uncorrelated Noise Addition
abstract
Personal data are of great interest in statistical studies and to provide personalized services, but its release may impair the privacy of individuals. To protect the privacy, in this paper, we present the notion and practical enforcement of semantic noise, a semantically-grounded version of the numerical uncorrelated noise addition method, which is capable of masking textual data while properly preserving their semantics. Unlike other perturbative masking schemes, our method can work with both datasets containing information of several individuals and single data. Empirical results show that our proposal provides semantically-coherent outcomes preserving data utility better than non-semantic perturbative mechanisms.
Mercedes Rodriguez-Garcia, Montserrat Batet, David Sánchez 0001
ICTAI3
2015 Ontology-Based Delegation of Access Control: An Enhancement to the XACML Delegation Profile
Malik Imran Daud, David Sánchez 0001, Alexandre Viejo
TrustBus2
2015 Semantic variance: An intuitive measure for ontology accuracy evaluation
David Sánchez 0001, Montserrat Batet, Sergio Martínez, Josep Domingo-Ferrer
Eng. Appl. Artif. Intell.1
2015 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation
abstract
Microaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate k-anonymous data sets, where the identity of each subject is hidden within a group of k subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers, and avoiding discretization of numerical data. k-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of k subjects is too small. To address this issue, several refinements of k-anonymity have been proposed, among which t-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate t-close data sets are based on generalization and suppression (they are extensions of k-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate k-anonymous t-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
IEEE Trans. Knowl. Data Eng.3
2014 Semantic Anonymisation of Set-valued Data
abstract
It is quite common that companies and organizations require of releasing and exchanging information related to individuals. Due to the usual sensitive nature of these data, appropriate measures should be applied to reduce the risk of re-identification of individuals while keeping as much data utility as possible. Many anonymization mechanisms have been developed up to present, even though most of them focus on structured/relational databases containing numerical or categorical data. However, the anonymization of transactional data, also known as set-valued data, has received much less attention. The management and transformation of these data presents additional challenges due to their variable cardinality and their usually textual and unbounded nature. Current approaches focusing on set-valued data are based on the generalization of original values; however, this suffers from a high information loss derived from the reduced granularity of output values. To tackle this problem, in this paper we adapt a well-known microaggregation anonymization mechanism so that it can be applied to set-valued data. Moreover, since the utility of textual data is closely related to their meaning, special care has been put in improving the preservation of data semantics. To do so, semantic similarity and aggregation functions are proposed. Experiments conducted on a real set-valued data set show that our proposal better preserves data utility in comparison with non-semantic approaches.
Montserrat Batet, Arnau Erola, David Sánchez 0001, Jordi Castellà-Roca
ICAART (1)3
2014 A Semantic Approach for Ontology Evaluation
abstract
In recent years, ontologies have experienced an enormous development due to their importance in knowledge-based systems. Because of the discrepancies that may appear during the modeling of ontologies, and due to the availability of ontologies covering overlapping domains, ontology evaluation is crucial in order to select the most appropriate ontology for a specific application. Many ontology evaluation mechanisms available in the literature assess the quality of ontologies according to their structural features. Even though, most of them propose ad-hoc scores aggregating different features, which lacks semantic and mathematical coherence. In this paper, we present an intuitive measure for ontology evaluation that quantifies the semantic dispersion of the ontology, which is both mathematically and semantically coherent. Our proposal is inspired in the standard notion of numerical dispersion of a sample and on a recent empirical study showing which ontological features can better predict the accuracy of ontologies. Our experiments, performed over a set of widely used ontologies, suggest that our measure positively correlates with such features, while offering a more coherent ontology evaluation score.
Montserrat Batet, David Sánchez 0001
ICTAI2
2014 Privacy protection of textual medical documents
abstract
With the adoption of ITs, a large amount patient-related documents is compiled by healthcare organisations. Quite often, this data is needed to be released to third parties for research or business purposes. The inherent sensitivity of patient's information has brought to the definition of legislations to protect the privacy of individuals. To meet with these legislations, redaction or sanitization of patient-related documents is needed before their release. This is usually done manually, which is costly and time-consuming, or by means of ad-hoc solutions that just protect structured types of sensitive information (e.g. social security numbers), or that are based on removing sensitive terms, which hampers the utility of the output. In this paper, we propose an automatic sanitization method for textual medical documents that is able to protect sensitive terms and those that are semantically related, while retaining the utility of the output as much as possible. Different to redaction schemas, which are based on term removal, our method improves the utility of the protected output by replacing sensitive terms with appropriate generalisations retrieved from medical and general-purpose knowledge bases. Experiments conducted on highly sensitive documents and in coherency with current regulations on healthcare data privacy show promising results in terms of output's privacy and utility.
Montserrat Batet, David Sánchez 0001
NOMS2
2014 Improving the Utility of Differential Privacy via Univariate Microaggregation
David Sánchez 0001, Josep Domingo-Ferrer, Sergio Martínez
Privacy in Statistical Databases1
2014 Distance Computation between Two Private Preference Functions
Alberto Blanco-Justicia, Josep Domingo-Ferrer, Oriol Farràs, David Sánchez 0001
SEC4
2014 Utility-preserving sanitization of semantically correlated terms in textual documents
David Sánchez 0001, Montserrat Batet, Alexandre Viejo
Inf. Sci.1
2014 An information theoretic approach to improve semantic similarity assessments across multiple ontologies
Montserrat Batet, Sébastien Harispe, Sylvie Ranwez, David Sánchez 0001, Vincent Ranwez
Inf. Sci.4
2014 Profiling social networks to provide useful and privacy-preserving web search
abstract
Web search engines (WSEs) use search queries to profile users and to provide personalized services like query disambiguation or refinement. These services are valuable because users get an enhanced search experience. However, the compiled user profiles may contain sensitive information that might represent a privacy threat. This issue should be addressed in a way that it also preserves the utility of the profile with regard to search services. State‐of‐the‐art approaches tackle these issues by generating and submitting fake queries that are related to the interests of the user. This technique allows the WSE to only know general (and useful) data while the detailed (and potentially private) data are obfuscated. To build fake queries, these proposals rely on past queries to obtain user interests. However, we argue that this is not always the best strategy and, in this article, we study the use of social networks to gather more accurate user profiles that enable better personalized service while offering a similar, or even better, level of practical privacy. These hypotheses are empirically supported by evaluations using real profiles gathered from Twitter and a set of AOL search queries.
Alexandre Viejo, David Sánchez 0001
J. Assoc. Inf. Sci. Technol.2
2014 Utility-preserving privacy protection of textual healthcare documents
David Sánchez 0001, Montserrat Batet, Alexandre Viejo
J. Biomed. Informatics1
2014 A framework for unifying ontology-based semantic similarity measures: A study in the biomedical domain
Sébastien Harispe, David Sánchez 0001, Sylvie Ranwez, Stefan Janaqi, Jacky Montmain
J. Biomed. Informatics2
2014 Towards the estimation of feature-based semantic similarity using multiple ontologies
Albert Solé-Ribalta, David Sánchez 0001, Montserrat Batet, Francesc Serratosa
Knowl. Based Syst.2
2014 Enhancing data utility in differential privacy via microaggregation-based k-anonymity
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
VLDB J.3
2013 Providing useful and private Web search by means of social network profiling
abstract
Web search engines (WSEs) build user profiles and use them to offer an enhanced web search experience. Nevertheless, these elements might contain sensitive data that may represent a privacy threat for the users. There are some works in the literature that address this situation while preserving the profile usefulness. These schemes submit synthetic queries that are fake but related to the real general interests of the user. Specifically, they rely on past user queries to obtain the legitimate interests of each user. We argue that this is not always the best strategy and, in this paper, we study the use of social networks to gather this information and provide a better personalized service while offering an equivalent privacy level.
Alexandre Viejo, David Sánchez 0001
PST2
2013 Semantic similarity estimation from multiple ontologies
Montserrat Batet, David Sánchez 0001, Aïda Valls, Karina Gibert
Appl. Intell.2
2013 An automatic approach for ontology-based feature extraction from heterogeneous textualresources
Carlos Vicient, David Sánchez 0001, Antonio Moreno
Eng. Appl. Artif. Intell.2
2013 A semantic similarity method based on information content exploiting multiple ontologies
David Sánchez 0001, Montserrat Batet
Expert Syst. Appl.1
2013 Minimizing the disclosure risk of semantic correlations in document sanitization
David Sánchez 0001, Montserrat Batet, Alexandre Viejo
Inf. Sci.1
2013 Utility preserving query log anonymization via semantic microaggregation
Montserrat Batet, Arnau Erola, David Sánchez 0001, Jordi Castellà-Roca
Inf. Sci.3
2013 Anonymization of nominal data based on semantic marginality
Josep Domingo-Ferrer, David Sánchez 0001, Guillem Rufian-Torrell
Inf. Sci.2
2013 Knowledge-based scheme to create privacy-preserving but semantically-related queries for web search engines
David Sánchez 0001, Jordi Castellà-Roca, Alexandre Viejo
Inf. Sci.1
2013 A semantic framework to protect the privacy of electronic health records with non-numerical attributes
Sergio Martínez, David Sánchez 0001, Aïda Valls
J. Biomed. Informatics2
2013 Automatic General-Purpose Sanitization of Textual Documents
abstract
The advent of new information sharing technologies has led society to a scenario where thousands of textual documents are publicly published every day. The existence of confidential information in many of these documents motivates the use of measures to hide sensitive data before being published, which is precisely the goal of document sanitization. Even though methods to assist the sanitization process have been proposed, most of them are focused on the detection of specific types of sensitive entities for concrete domains, lacking generality and and requiring user supervision. Moreover, to hide sensitive terms, most approaches opt to remove them, a measure that hampers the utility of the sanitized document. This paper presents a general-purpose sanitization method that, based on information theory and exploiting knowledge bases, detects and hides sensitive textual information while preserving its meaning. Our proposal works in an automatic and unsupervised way and it can be applied to heterogeneous documents, which make it specially suitable for environments with massive and heterogeneous information-sharing needs. Evaluation results show that our method outperforms strategies based on trained classifiers regarding the detection recall, whereas it better retains the document's utility compared to term-suppression methods.
David Sánchez 0001, Montserrat Batet, Alexandre Viejo
IEEE Trans. Inf. Forensics Secur.1
2012 Towards k-Anonymous Non-numerical Data via Semantic Resampling
Sergio Martínez, David Sánchez 0001, Aïda Valls
IPMU (4)2
2012 Detecting Sensitive Information from Textual Documents: An Information-Theoretic Approach
David Sánchez 0001, Montserrat Batet, Alexandre Viejo
MDAI1
2012 Using Profiling Techniques to Protect the User's Privacy in Twitter
Alexandre Viejo, David Sánchez 0001, Jordi Castellà-Roca
MDAI2
2012 Semantic adaptive microaggregation of categorical microdata
Sergio Martínez, David Sánchez 0001, Aïda Valls
Comput. Secur.2
2012 Turist@: Agent-based personalised recommendation of tourist activities
Montserrat Batet, Antonio Moreno, David Sánchez 0001, David Isern, Aïda Valls
Expert Syst. Appl.3
2012 Ontology-based semantic similarity: A new feature-based approach
David Sánchez 0001, Montserrat Batet, David Isern, Aïda Valls
Expert Syst. Appl.1
2012 Learning relation axioms from text: An automatic Web-based approach
David Sánchez 0001, Antonio Moreno, Luis Del Vasto-Terrientes
Expert Syst. Appl.1
2012 A New Model to Compute the Information Content of Concepts from Taxonomic Knowledge
abstract
The Information Content (IC) of a concept quantifies the amount of information it provides when appearing in a context. In the past, IC used to be computed as a function of concept appearance probabilities in corpora, but corpora-dependency and data sparseness hampered results. Recently, some other authors tried to overcome previous approaches, estimating IC from the knowledge modeled in an ontology. In this paper, the authors develop this idea, by proposing a new model to compute the IC of a concept exploiting the taxonomic knowledge modeled in an ontology. In comparison with related works, their proposal aims to better capture semantic evidences found in the ontology. To test the authors’ approach, they have applied it to well-known semantic similarity measures, which were evaluated using standard benchmarks. Results show that the use of the authors’ model produces, in most cases, more accurate similarity estimations than related works.
David Sánchez 0001, Montserrat Batet
Int. J. Semantic Web Inf. Syst.1
2012 Enabling semantic similarity estimation across multiple ontologies: An evaluation in the biomedical domain
David Sánchez 0001, Albert Solé-Ribalta, Montserrat Batet, Francesc Serratosa
J. Biomed. Informatics1
2012 Knowledge-driven delivery of home care services
Montserrat Batet, David Isern, Lucas Marin, Sergio Martínez, Antonio Moreno, David Sánchez 0001, Aïda Valls, Karina Gibert
J. Intell. Inf. Syst.6
2012 Semantically-grounded construction of centroids for datasets with textual attributes
Sergio Martínez, Aïda Valls, David Sánchez 0001
Knowl. Based Syst.3
2012 Preventing automatic user profiling in Web 2.0 applications
Alexandre Viejo, David Sánchez 0001, Jordi Castellà-Roca
Knowl. Based Syst.2
2011 Semantic-based Composition of Modular Ontologies Applied to Web Query Reformulation
Manel Elloumi-Chaabene, Nesrine Ben Mustapha, Hajer Baazaoui Zghal, Antonio Moreno, David Sánchez 0001
ICSOFT (1)5
2011 Agent-based execution of personalised home care treatments
David Isern, Antonio Moreno, David Sánchez 0001, Ákos Hajnal, Gianfranco Pedone, László Z. Varga
Appl. Intell.3
2011 Automatic extraction of acronym definitions from the Web
David Sánchez 0001, David Isern
Appl. Intell.1
2011 Agent-based platform to support the execution of parallel tasks
David Sánchez 0001, David Isern, Ángel Rodríguez-Rozas, Antonio Moreno
Expert Syst. Appl.1
2011 An ontology-based measure to compute semantic similarity in biomedicine
Montserrat Batet, David Sánchez 0001, Aïda Valls
J. Biomed. Informatics2
2011 Semantic similarity estimation in the biomedical domain: An ontology-based information-theoretic perspective
David Sánchez 0001, Montserrat Batet
J. Biomed. Informatics1
2011 Organizational structures supported by agent-oriented methodologies
David Isern, David Sánchez 0001, Antonio Moreno
J. Syst. Softw.2
2011 Content annotation for the semantic web: an automatic web-based approach
David Sánchez 0001, David Isern, Miquel Millan
Knowl. Inf. Syst.1
2011 Ontology-based information content computation
David Sánchez 0001, Montserrat Batet, David Isern
Knowl. Based Syst.1
2010 Semantic Web Search System Founded on Case-Based Reasoning and Ontology Learning
Hajer Baazaoui Zghal, Nesrine Ben Mustapha, Manel Elloumi-Chaabene, Antonio Moreno, David Sánchez 0001
IC3K5
2010 Exploiting Taxonomical Knowledge to Compute Semantic Similarity: An Evaluation in the Biomedical Domain
Montserrat Batet, David Sánchez 0001, Aïda Valls, Karina Gibert
IEA/AIE (1)2
2010 Anonymizing Categorical Data with a Recoding Method Based on Semantic Similarity
Sergio Martínez, Aïda Valls, David Sánchez 0001
IPMU (2)3
2010 Discovery of Relation Axioms from the Web
Luis Del Vasto-Terrientes, Antonio Moreno, David Sánchez 0001
KSEM3
2010 Ontology-Based Anonymization of Categorical Values
Sergio Martínez, David Sánchez 0001, Aïda Valls
MDAI2
2010 A methodology to learn ontological attributes from the Web
David Sánchez 0001
Data Knowl. Eng.1
2010 Ontology-driven web-based semantic similarity
David Sánchez 0001, Montserrat Batet, Aïda Valls, Karina Gibert
J. Intell. Inf. Syst.1
2009 Computing Knowledge-Based Semantic Similarity from the Web: An Application to the Biomedical Domain
David Sánchez 0001, Montserrat Batet, Aïda Valls
KSEM1
2008 Learning non-taxonomic relationships from web documents for domain ontology construction
David Sánchez 0001, Antonio Moreno
Data Knowl. Eng.1
2007 An Ontology-Driven Agent-Based Clinical Guideline Execution Engine
David Isern, David Sánchez 0001, Antonio Moreno
AIME2
2006 Discovering Non-taxonomic Relations from the Web
David Sánchez 0001, Antonio Moreno
IDEAL1
2006 Integrated Agent-Based Approach for Ontology-Driven Web Filtering
David Sánchez 0001, David Isern, Antonio Moreno
KES (3)1
2005 Web Mining Techniques for Automatic Discovery of Medical Knowledge
David Sánchez 0001, Antonio Moreno
AIME1
2005 Development of new techniques to improve Web search
David Sánchez 0001, Antonio Moreno
IJCAI1