Josep Domingo-Ferrer

dblp:d/JDomingoFerrer · DBLP profile ↗
← Back
50ranked-venue papers in the field
25as first author
7since 2021 · last 2025
0000-0001-7213-4962ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 20 (11 first)Database Systems & Data Management · 16 (8 first)Data Mining & Knowledge Discovery · 9 (3 first)Other / Interdisciplinary · 5 (3 first)
YearPublicationVenuePosition
2025 Enhancing efficiency and data utility in longitudinal data anonymization
Fatemeh Amiri, David Sánchez 0001, Josep Domingo-Ferrer
Inf. Sci.3
2024 Machine Learning in Finance
abstract
This workshop aims to explore the intersection of Generative AI with the rich tapestry of financial data types, seeking to uncover new methodologies and techniques that can enhance predictive analytics, fraud detection, and customer insights across the sector. By harnessing these advancements in AI, we can pave the way to not only understand customer behavior but also anticipate their needs more effectively, leading to superior customer outcomes and more personalized services. Our objective is to shed light on the challenges and opportunities presented by the diverse data formats in finance. We aim to bridge the gap between the dominance of traditional models for tabular data analysis and the emerging potential of Generative AI to revolutionize the treatment of time series, click streams, and other unstructured data forms.
Leman Akoglu, Nitesh V. Chawla, Josep Domingo-Ferrer, Eren Kurshan, Senthil Kumar, Vidyut M. Naware, José A. Rodríguez-Serrano, Isha Chaturvedi, Saurabh Nagrecha, Mahashweta Das, Tanveer A. Faruquie
KDD3
2023 Defending Against Backdoor Attacks by Layer-wise Feature Analysis
Najeeb Jebreel, Josep Domingo-Ferrer, Yiming Li 0004
PAKDD (2)2
2023 Fair detection of poisoning attacks in federated learning on non-i.i.d. data
Ashneet Khandpur Singh, Alberto Blanco-Justicia, Josep Domingo-Ferrer
Data Min. Knowl. Discov.3
2023 Utility-Preserving Privacy Protection of Textual Documents via Word Embeddings
abstract
A great variety of mechanisms have been proposed to protect structured databases with numerical and categorical attributes; however, little attention has been devoted to unstructured textual data. Textual data protection requires first detecting sensitive pieces of text and then masking those pieces via suppression or generalization. Current solutions rely on classifiers that can recognize a fixed set of (allegedly sensitive) named entities. Yet, such approaches fall short of providing adequate protection because in reality references to sensitive information are not limited to a predefined set of entity types, and not all the appearances of certain entity type result in disclosure. In this work we propose a more general and flexible based on the notion of word embedding. By means of word embeddings we build vectors that numerically capture the semantic relationships of the textual terms. Then we evaluate the disclosure caused by the terms on the entity to be protected according to the similarity between their vector representations. Our method also preserves the semantics (and, therefore, the utility) of the document by replacing risky terms with privacy-preserving generalizations. Empirical results show that our approach offers much more robust protection and greater utility preservation than methods based on named entity recognition.
Fadi Hassan, David Sánchez 0001, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.3
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
ICDE1
2022 Multi-Dimensional Randomized Response
abstract
In our data world, a host of not necessarily trusted controllers gather data on individual subjects. To preserve her privacy and, more generally, her informational self-determination, the individual has to be empowered by giving her agency on her own data. Maximum agency is afforded by local anonymization, that allows each individual to anonymize her own data before handing them to the data controller. Randomized response (RR) is a local anonymization approach able to yield multi-dimensional full sets of anonymized microdata that are valid for exploratory analysis and machine learning. This is so because an unbiased estimate of the distribution of the true data of individuals can be obtained from their pooled randomized data. Furthermore, RR offers rigorous privacy guarantees. The main weakness of RR is the curse of dimensionality when applied to several attributes: as the number of attributes grows, the accuracy of the estimated true data distribution quickly degrades. We propose several complementary approaches to mitigate the dimensionality problem. First, we present two basic protocols, separate RR on each attribute and joint RR for all attributes, and discuss their limitations. Then we introduce an algorithm to form clusters of attributes so that attributes in different clusters can be viewed as independent and joint RR can be performed within each cluster. After that, we introduce an adjustment algorithm for the randomized data set that repairs some of the accuracy loss due to assuming independence between attributes when using RR separately on each attribute or due to assuming independence between clusters in cluster-wise RR. We also present empirical work to illustrate the proposed methods.
Josep Domingo-Ferrer, Jordi Soria-Comas
IEEE Trans. Knowl. Data Eng.1
2020 Outsourcing analyses on privacy-protected multivariate categorical data stored in untrusted clouds
Josep Domingo-Ferrer, David Sánchez 0001, Sara Ricci, Mónica Muñoz-Batista
Knowl. Inf. Syst.1
2019 Personal Big Data, GDPR and Anonymization
Josep Domingo-Ferrer
FQAS1
2018 Outsourcing scalar products and matrix products on privacy-protected unencrypted data stored in untrusted clouds
Josep Domingo-Ferrer, Sara Ricci, Carles Domingo-Enrich
Inf. Sci.1
2018 Co-utile disclosure of private data in social networks
David Sánchez 0001, Josep Domingo-Ferrer, Sergio Martínez
Inf. Sci.2
2016 t-closeness through microaggregation: Strict privacy with enhanced utility preservation
abstract
This paper proposes and shows how to use microaggregation to attain t-closeness on top of k-anonymity to protect data releases. The advantages in terms of data utility preservation of microaggregation over classic approaches based on generalizing values are analyzed. Then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
ICDE2
2016 Big Data Privacy: Challenges to Privacy Principles and Models
abstract
Abstract This paper explores the challenges raised by big data in privacy-preserving data management. First, we examine the conflicts raised by big data with respect to preexisting concepts of private data management, such as consent, purpose limitation, transparency and individual rights of access, rectification and erasure. Anonymization appears as the best tool to mitigate such conflicts, and it is best implemented by adhering to a privacy model with precise privacy guarantees. For this reason, we evaluate how well the two main privacy models used in anonymization (k-anonymity and $$\varepsilon $$ ε -differential privacy) meet the requirements of big data, namely composability, low computational cost and linkability.
Jordi Soria-Comas, Josep Domingo-Ferrer
Data Sci. Eng.2
2016 Self-enforcing protocols via co-utile reputation management
Josep Domingo-Ferrer, Oriol Farràs, Sergio Martínez, David Sánchez 0001, Jordi Soria-Comas
Inf. Sci.1
2016 New directions in anonymization: Permutation paradigm, verifiability by subjects and intruders, transparency to users
Josep Domingo-Ferrer, Krishnamurty Muralidhar
Inf. Sci.1
2015 Discrimination- and privacy-aware patterns
Sara Hajian, Josep Domingo-Ferrer, Anna Monreale, Dino Pedreschi, Fosca Giannotti
Data Min. Knowl. Discov.2
2015 t-Closeness through Microaggregation: Strict Privacy with Enhanced Utility Preservation
abstract
Microaggregation is a technique for disclosure limitation aimed at protecting the privacy of data subjects in microdata releases. It has been used as an alternative to generalization and suppression to generate k-anonymous data sets, where the identity of each subject is hidden within a group of k subjects. Unlike generalization, microaggregation perturbs the data and this additional masking freedom allows improving data utility in several ways, such as increasing data granularity, reducing the impact of outliers, and avoiding discretization of numerical data. k-Anonymity, on the other side, does not protect against attribute disclosure, which occurs if the variability of the confidential values in a group of k subjects is too small. To address this issue, several refinements of k-anonymity have been proposed, among which t-closeness stands out as providing one of the strictest privacy guarantees. Existing algorithms to generate t-close data sets are based on generalization and suppression (they are extensions of k-anonymization algorithms based on the same principles). This paper proposes and shows how to use microaggregation to generate k-anonymous t-close data sets. The advantages of microaggregation are analyzed, and then several microaggregation algorithms for k-anonymous t-closeness are presented and empirically evaluated.
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
IEEE Trans. Knowl. Data Eng.2
2014 Generalization-based privacy preservation and discrimination prevention in data publishing and mining
Sara Hajian, Josep Domingo-Ferrer, Oriol Farràs
Data Min. Knowl. Discov.2
2014 Ciphertext-policy hierarchical attribute-based encryption with short ciphertexts
Qianhong Wu, Josep Domingo-Ferrer, Lei Zhang 0009, Jianwei Liu 0001, Wenchang Shi
Inf. Sci.4
2014 Signatures in hierarchical certificateless cryptography: Efficient constructions and provable security
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.3
2014 Enhancing data utility in differential privacy via microaggregation-based k-anonymity
Jordi Soria-Comas, Josep Domingo-Ferrer, David Sánchez 0001, Sergio Martínez
VLDB J.2
2013 On the privacy offered by (k, δ)-anonymity
Rolando Trujillo-Rasua, Josep Domingo-Ferrer
Inf. Syst.2
2013 Anonymization of nominal data based on semantic marginality
Josep Domingo-Ferrer, David Sánchez 0001, Guillem Rufian-Torrell
Inf. Sci.1
2013 Optimal data-independent noise for differential privacy
Jordi Soria-Comas, Josep Domingo-Ferrer
Inf. Sci.2
2013 A Methodology for Direct and Indirect Discrimination Prevention in Data Mining
abstract
Data mining is an increasingly important technology for extracting useful knowledge hidden in large collections of data. There are, however, negative social perceptions about data mining, among which potential privacy invasion and potential discrimination. The latter consists of unfairly treating people on the basis of their belonging to a specific group. Automated data collection and data mining techniques such as classification rule mining have paved the way to making automated decisions, like loan granting/denial, insurance premium computation, etc. If the training data sets are biased in what regards discriminatory (sensitive) attributes like gender, race, religion, etc., discriminatory decisions may ensue. For this reason, anti-discrimination techniques including discrimination discovery and prevention have been introduced in data mining. Discrimination can be either direct or indirect. Direct discrimination occurs when decisions are made based on sensitive attributes. Indirect discrimination occurs when decisions are made based on nonsensitive attributes which are strongly correlated with biased sensitive ones. In this paper, we tackle discrimination prevention in data mining and propose new techniques applicable for direct or indirect discrimination prevention individually or both at the same time. We discuss how to clean training data sets and outsourced data sets in such a way that direct and/or indirect discriminatory decision rules are converted to legitimate (nondiscriminatory) classification rules. We also propose new metrics to evaluate the utility of the proposed approaches and we compare these approaches. The experimental evaluations demonstrate that the proposed techniques are effective at removing direct and/or indirect discrimination biases in the original data set while preserving data quality.
Sara Hajian, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.2
2012 Rational behavior in peer-to-peer profile obfuscation for anonymous keyword search
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2012 Rational behavior in peer-to-peer profile obfuscation for anonymous keyword search: The multi-hop scenario
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2012 Microaggregation- and permutation-based anonymization of movement data
Josep Domingo-Ferrer, Rolando Trujillo-Rasua
Inf. Sci.1
2012 Provably secure threshold public-key encryption with adaptive security and short ciphertexts
Qianhong Wu, Lei Zhang 0009, Oriol Farràs, Josep Domingo-Ferrer
Inf. Sci.5
2011 Provably secure one-round identity-based authenticated asymmetric group key agreement protocol
Lei Zhang 0009, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.4
2010 Hybrid microdata using microaggregation
Josep Domingo-Ferrer, Úrsula González-Nicolás
Inf. Sci.1
2010 Simulatable certificateless two-party authenticated key agreement protocol
Lei Zhang 0009, Futai Zhang, Qianhong Wu, Josep Domingo-Ferrer
Inf. Sci.4
2010 From t-Closeness-Like Privacy to Postrandomization via Information Theory
abstract
t-Closeness is a privacy model recently defined for data anonymization. A data set is said to satisfy t-closeness if, for each group of records sharing a combination of key attributes, the distance between the distribution of a confidential attribute in the group and the distribution of the attribute in the entire data set is no more than a threshold t. Here, we define a privacy measure in terms of information theory, similar to t-closeness. Then, we use the tools of that theory to show that our privacy measure can be achieved by the postrandomization method (PRAM) for masking in the discrete case, and by a form of noise addition in the general case.
David Rebollo-Monedero, Jordi Forné, Josep Domingo-Ferrer
IEEE Trans. Knowl. Data Eng.3
2009 User-private information retrieval based on a peer-to-peer community
Josep Domingo-Ferrer, Maria Bras-Amorós, Qianhong Wu, Jesús A. Manjón
Data Knowl. Eng.1
2009 Recent progress in database privacy
Josep Domingo-Ferrer, Yücel Saygin
Data Knowl. Eng.1
2009 Erratum to "A measure of variance for hierarchical nominal attributes"
Josep Domingo-Ferrer, Agusti Solanas
Inf. Sci.1
2008 A measure of variance for hierarchical nominal attributes
Josep Domingo-Ferrer, Agusti Solanas
Inf. Sci.1
2008 Efficient Remote Data Possession Checking in Critical Information Infrastructures
abstract
Checking data possession in networked information systems such as those related to critical infrastructures (power facilities, airports, data vaults, defense systems, etc.) is a matter of crucial importance. Remote data possession checking protocols permit to check that a remote server can access an uncorrupted file in such a way that the verifier does not need to know beforehand the entire file that is being verified. Unfortunately, current protocols only allow a limited number of successive verifications or are impractical from the computational point of view. In this paper, we present a new remote data possession checking protocol such that: i) it allows an unlimited number of file integrity verifications; ii) its maximum running time can be chosen at set-up time and traded off against storage at the verifier.
Francesc Sebé, Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Yves Deswarte, Jean-Jacques Quisquater
IEEE Trans. Knowl. Data Eng.2
2006 Regression for ordinal variables without underlying continuous variables
Vicenç Torra, Josep Domingo-Ferrer, Josep Maria Mateo-Sanz, Michael Kwok-Po Ng
Inf. Sci.2
2006 Efficient multivariate data-oriented microaggregation
Josep Domingo-Ferrer, Antoni Martínez-Ballesté, Josep Maria Mateo-Sanz, Francesc Sebé
VLDB J.1
2005 Privacy in Data Mining
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.1
2005 Ordinal, Continuous and Heterogeneous k-Anonymity Through Microaggregation
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.1
2005 Probabilistic Information Loss Measures in Confidentiality Protection of Continuous Microdata
Josep Maria Mateo-Sanz, Josep Domingo-Ferrer, Francesc Sebé
Data Min. Knowl. Discov.2
2003 Median-based aggregation operators for prototype construction in ordinal scales
abstract
This article studies aggregation operators in ordinal scales for their application to clustering (more specifically, to microaggregation for statistical disclosure risk). In particular, we consider these operators in the process of prototype construction. This study analyzes main aggregation operators for ordinal scales [plurality rule, medians, Sugeno integrals (SI), and ordinal weighted means (OWM), among others] and shows the difficulties for their application in this particular setting. Then, we propose two approaches to solve the drawbacks and we study their properties. Special emphasis is given to the study of monotonicity because the operator is proven nonsatisfactory for this property. Exhaustive empirical work shows that in most practical situations, this cannot be considered a problem. © 2003 Wiley Periodicals, Inc.
Josep Domingo-Ferrer, Vicenç Torra
Int. J. Intell. Syst.1
2003 Semantic-based aggregation for statistical disclosure control
abstract
In this paper we show how clustering can be used to aggregate different versions of the same data set in order to discover confidential information. Having these tools helps to not publish data that could be reidentified, which is known as Statistical Disclosure Control. In particular, the paper is focused on the case of dealing with categorical values. © 2003 Wiley Periodicals, Inc.
Aïda Valls, Vicenç Torra, Josep Domingo-Ferrer
Int. J. Intell. Syst.3
2003 On the connections between statistical disclosure control for microdata and some artificial intelligence tools
Josep Domingo-Ferrer, Vicenç Torra
Inf. Sci.1
2002 Information-Theoretic Disclosure Risk Measures in Statistical Disclosure Control of Tabular Data
abstract
Statistical database protection is a part of information security which tries to prevent published statistical information (tables, individual records) from disclosing the contribution of specific respondents. This paper shows how to use information-theoretic concepts to measure disclosure risk for tabular data. The proposed disclosure risk measure is compatible with a broad class of disclosure protection methods and can be extended for computing disclosure risk for a set of linked tables.
Josep Domingo-Ferrer, Anna Oganian, Vicenç Torra
SSDBM1
2002 Practical Data-Oriented Microaggregation for Statistical Disclosure Control
abstract
Microaggregation is a statistical disclosure control technique for microdata disseminated in statistical databases. Raw microdata (i.e., individual records or data vectors) are grouped into small aggregates prior to publication. Each aggregate should contain at least k data vectors to prevent disclosure of individual information, where k is a constant value preset by the data protector. No exact polynomial algorithms are known to date to microaggregate optimally, i.e., with minimal variability loss. Methods in the literature rank data and partition them into groups of fixed-size; in the multivariate case, ranking is performed by projecting data vectors onto a single axis. In this paper, candidate optimal solutions to the multivariate and univariate microaggregation problems are characterized. In the univariate case, two heuristics based on hierarchical clustering and genetic algorithms are introduced which are data-oriented in that they try to preserve natural data aggregates. In the multivariate case, fixed-size and hierarchical clustering microaggregation algorithms are presented which do not require data to be projected onto a single dimension; such methods clearly reduce variability loss as compared to conventional multivariate microaggregation on projected data.
Josep Domingo-Ferrer, Josep Maria Mateo-Sanz
IEEE Trans. Knowl. Data Eng.1
1996 A New Privacy Homomorphism and Applications
Josep Domingo-Ferrer
Inf. Process. Lett.1
1991 Distributed User Identification by Zero-Knowledge Access Rights Proving
Josep Domingo-Ferrer
Inf. Process. Lett.1