Vicenç Torra

dblp:t/VicencTorra · also Vicenç Torra I. Reventós · DBLP profile ↗
← Back
209ranked-venue papers
84as first author
41since 2021 · last 2026
0000-0002-0368-8037ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 147 · 69 first-author · 22 since 2021Databases, data management, data science and information retrieval · 53 · 25 first-author · 6 since 2021Security and privacy · 37 · 7 first-author · 16 since 2021Theory of computation · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorComputer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Approximating fuzzy measures with distorted probabilities
abstract
Fuzzy measures, also called capacities, non-additive measures, and monotonic games, are an interesting mathematical object and have been used in a large number of real-world applications. They are useful in problems where we need to aggregate information. In this case, we often use them in combination with fuzzy integrals. Sugeno and Choquet integrals are classical examples of fuzzy integrals. Fuzzy measures are set functions, and as such, when the reference set is finite, they are defined by 2 n values where n represents the cardinality of the reference set. This makes difficult both to define them and to interpret them. In this work, we propose to approximate fuzzy measures in terms of distorted probabilities as a way to better understand the structure and properties of the measure. This approach can be used when fuzzy measures are learned from examples. We show, however, that, in general, the approximation will not represent all the properties of an arbitrary fuzzy measure.
Vicenç Torra
Fuzzy Sets Syst.1
2026 Set contribution functions for quantitative bipolar argumentation and their principles
Filip Naudot, Andreas Brännström, Vicenç Torra, Timotheus Kampik
Int. J. Approx. Reason.3
2026 PrunePrivyTune: Accelerating Language Models with Pruning and Differentially Private Fine-Tuning
abstract
Abstract Large Language Models (LLMs) have demonstrated exceptional capabilities in language understanding and generation, but their large-scale architecture poses significant challenges in deployment and inference, such as increased computational demands and slower processing times. While various techniques like model pruning, knowledge distillation, and quantization have been developed to compress LLMs, they often result in task-specific compression, limiting the model’s versatility. Additionally, LLMs face privacy risks due to their potential to memorize and reproduce sensitive training data, raising concerns when deployed in real-world applications. To address these challenges, we propose a novel methodology PrunePrivyTune that combines efficient model compression with privacy preserving fine-tuning. Our approach leverages pairwise cosine similarity to identify redundant layers in transformer models, enabling structural pruning that reduces model size without compromising performance. After pruning, we apply Low-Rank Adaptation (LoRA) with DPSGD to fine-tune the model. This ensures that fine-tuning process is both efficient and privacy-preserving, outperforming training and preventing the model from memorizing sensitive data. Later on, we generated synthetic data using the fine-tuned model and subsequently conducted a training data extraction attack to assess the model’s privacy vulnerabilities, in terms of perplexity and BERTScore. Our framework demonstrates that the proposed methodology effectively reduces the inference time through model compression and pruning compliments privacy, followed by private fine-tuning. Additionally, our privacy risk assessment indicates that integrating DP successfully mitigates the risk of the model’s memorization. This approach upholds strong privacy guarantees, making it highly suitable for real-time applications and deployment in sensitive domains where data confidentiality is paramount.
Sonakshi Garg, Vicenç Torra
Mach. Learn.2
2026 Fuzzy clustering-based microaggregation for multi-view data with constraints
abstract
Abstract Microaggregation is a powerful technique for safeguarding data, enabling us to strike a balance between the risk of disclosing sensitive information and the loss of valuable information. It is a crucial tool for data sharing that provides k -anonymity. With the growing prevalence of multi-view data, there is an increasing interest in protecting such data using appropriate techniques. To the best of our knowledge, this paper introduces the first approach specifically designed for multi-view data protection. We present a novel approach to microaggregation by introducing multi-view fuzzy c-means, which allows us to consider linear constraints on the variables in each view that describe the data. Our method ensures that the resulting clusters adhere to these constraints, even when the data being masked fails to satisfy them. This approach not only enhances data privacy by maintaining k -anonymity in multi-view contexts but also preserves the structural integrity of the data across different views. Our contributions include the development of a multi-view clustering framework with built-in privacy safeguards and the introduction of linear constraints to ensure consistency across multiple views. This innovative approach provides a robust solution for the privacy-preserving analysis and sharing of multi-view data.
Fatemeh Sadjadi, Vicenç Torra
Soft Comput.2
2025 Blockchain-Enhanced User Consent for GDPR-Compliant Real-Time Bidding
Cristòfol Daudén-Esmel, Jordi Castellà-Roca, Alexandre Viejo, Vicenç Torra
DBSec4
2025 A Study of the Fuzzy Differential Entropy
Zuzana Ontkovicová, Vicenç Torra
EUSFLAT (1)2
2025 Improving Locally Differentially Private Graph Statistics Through Sparseness-Preserving Noise-Graph Addition
abstract
Differential privacy allows to publish graph statistics in a way that protects individual privacy while stillallowing meaningful insights to be derived from the data. The centralized privacy model of differential privacyassumes that there is a trusted data curator, while the local model does not require such a trusted authority.Local differential privacy is commonly achieved through randomized response (RR) mechanisms. This doesnot preserve the sparseness of the graphs. As most of the real-world graphs are sparse and have several nodes,this is a drawback of RR-based mechanisms, in terms of computational efficiency and accuracy. We thus,propose a comparative analysis through experimental analysis and discussion, to compute statistics with localdifferential privacy, where, it is shown that preserving the sparseness of the original graphs is the key factorto gain that balance between utility and privacy. We perform several experiments to test the utility of theprotected graphs in terms of several sub-graph counting i.e. triangle, and star counting and other statistics. Weshow that the sparseness preserving algorithm gives comparable or better results in comparison to the otherstate of the art methods and improves computational efficiency.
Sudipta Paul 0008, Julián Salas, Vicenç Torra
ICISSP (2)3
2025 Privacy-Enhancing Federated Time-Series Forecasting: A Microaggregation-Based Approach
Sargam Gupta, Vicenç Torra
SECRYPT2
2025 Generalized F-spaces through the lens of fuzzy measures
abstract
Probabilistic metric spaces are natural extensions of metric spaces, where the function that computes the distance outputs a distribution on the real numbers rather than a single value. Such a function is called a distribution function. F-spaces are constructions for probabilistic metric spaces, where the distribution functions are built for functions that map from a measurable space to a metric space. In this paper, we propose an extension of F-spaces, called Generalized F-space. This construction replaces the metric space with a probabilistic metric space and uses fuzzy measures to evaluate sets of elements whose distances are probability distributions. We present several results that establish connections between the properties of the constructed space and specific fuzzy measures under particular triangular norms. Furthermore, we demonstrate how the space can be applied in machine learning to compute distances between different classifier models. Experimental results based on Sugeno λ -measures are consistent with our theoretical findings.
Mariam Taha, Vicenç Torra
Fuzzy Sets Syst.2
2025 Efficient federated unlearning under plausible deniability
abstract
Abstract Privacy regulations like the GDPR in Europe and the CCPA in the US allow users the right to remove their data from machine learning (ML) applications. Machine unlearning addresses this by modifying the ML parameters in order to forget the influence of a specific data point on its weights. Recent literature has highlighted that the contribution from data point(s) can be forged with some other data points in the dataset with probability close to one. This allows a server to falsely claim unlearning without actually modifying the model’s parameters. However, in distributed paradigms such as federated learning (FL), where the server lacks access to the dataset and the number of clients are limited, claiming unlearning in such cases becomes a challenge. An honest server must modify the model parameters in order to unlearn. This paper introduces an efficient way to achieve machine unlearning in FL, i.e., federated unlearning, by employing a privacy model which allows the FL server to plausibly deny the client’s participation in the training up to a certain extent. Specifically, we demonstrate that the server can generate a Proof-of-Deniability, where each aggregated update can be associated with at least x (the plausible deniability parameter) client updates. This enables the server to plausibly deny a client’s participation. However, in the event of frequent unlearning requests, the server is required to adopt an unlearning strategy and, accordingly, update its model parameters. We also perturb the client updates in a cluster in order to avoid inference from an honest but curious server. We show that the global model satisfies $$(\epsilon , \delta )$$ ( ϵ , δ ) -differential privacy after T number of communication rounds. The proposed methodology has been evaluated on multiple datasets in different privacy settings. The experimental results show that our framework achieves comparable utility while providing a significant reduction in terms of memory ( $$\approx $$ ≈ 30 times), as well as retraining time (1.6-500769 times). The source code for the paper is available https://github.com/Ayush-Umu/Federated-Unlearning-under-Plausible-Deniability .
Ayush K. Varshney, Vicenç Torra
Mach. Learn.2
2025 Unlearning Clients, Features and Samples in Vertical Federated Learning
abstract
Federated Learning ( FL ) has emerged as a prominent distributed learning paradigm that allows multiple users to collaboratively train a model without sharing their data thus preserving privacy. Within the scope of privacy preservation, information privacy regulations such as GDPR entitle users to request the removal (or unlearning) of their contribution from a service that is hosting the model. For this purpose, a server hosting an ML model must be able to unlearn certain information in cases such as copyright infringement or security issues that can make the model vulnerable or impact the performance of a service based on that model. While most unlearning approaches in FL focus on Horizontal Federated Learning (HFL), where clients share the feature space and the global model, Vertical Federated Learning (VFL) has received less attention from the research community. VFL involves clients (passive parties) sharing the sample space among them while not having access to the labels. In this paper, we explore unlearning in VFL from three perspectives: unlearning passive parties, unlearning features, and unlearning samples. To unlearn passive parties and features we introduce VFU-KD which is based on knowledge distillation (KD) while to unlearn samples, VFU-GA is introduced which is based on gradient ascent (GA). To provide evidence of approximate unlearning, we utilize Membership Inference Attack (MIA) to audit the effectiveness of our unlearning approach. Our experiments across six tabular datasets and two image datasets demonstrate that VFU-KD and VFU-GA achieve performance comparable to or better than both retraining from scratch and the benchmark R2S method in many cases, with improvements of (0 − 2%). In the remaining cases, utility scores remain comparable, with a modest utility loss ranging from 1 − 5%. Unlike existing methods, VFU-KD and VFU-GA require no communication between active and passive parties during unlearning. However, they do require the active party to store the previously communicated embeddings.
Ayush K. Varshney, Konstantinos Vandikas, Vicenç Torra
Proc. Priv. Enhancing Technol.3
2025 Uncoordinated Syntactic Privacy: A New Composable Metric for Multiple, Independent Data Publishing
abstract
A privacy model is a privacy condition, dependent on a parameter, that guarantees an upper bound on the risk of reidentification disclosure and maybe also on the risk of attribute disclosure by an adversary. A privacy model is composable if the privacy guarantees of the model are preserved, possibly to a limited extent, after repeated independent application of the privacy model. From the opposite perspective, a privacy model is not composable if multiple independent data releases, each of them satisfying the requirements of the privacy model, may result in a privacy breach. Current privacy models are broadly classified into syntactic ones (such as k-anonymity and l-diversity) and semantic ones, which essentially refer to$\varepsilon $-differential privacy (e-DP) and variations thereof. While e-DP and its variants offer strong composability properties, syntactic notions are not composable unless data releases are conducted by a single, centralized data holder that uses specialized notions such as m-invariance and$\tau $-safety. In this work, we propose m-uncoordinated-syntactic-privacy (m-USP), the first syntactic notion with composability properties for the independent publication of nondisjoint data, in other words, without a centralized data holder. Theoretical results are formally proven, and experimental results demonstrate that the risk to individuals does not increase significantly, in contrast to non-composable methods, that are susceptible to attribute disclosure. In most cases, the utility degradation caused by the extra protection is less than 5% and decreases as the value of m increases.
Adrián Tobar Nicolau, Javier Parra-Arnau, Jordi Forné, Vicenç Torra
IEEE Trans. Inf. Forensics Secur.4
2024 Task-Specific Knowledge Distillation with Differential Privacy in LLMs
Sonakshi Garg, Vicenç Torra
ESORICS (2)2
2024 The Past and Future of Fuzzy Measures and Fuzzy Integrals - -In Memory of Prof. Michio Sugeno
Yasuo Narukawa, Katsushige Fujimoto, Vicenç Torra, Zuzana Ontkovicová
MDAI3
2024 DISCOLEAF: Personalized DIScretization of COntinuous Attributes for LEArning with Federated Decision Trees
Saloni Kwatra, Vicenç Torra
PSD2
2024 Attribute Disclosure Risk in Smart Meter Data
Guillermo Navarro-Arribas, Vicenç Torra
PSD2
2024 Can Synthetic Data Preserve Manifold Properties?
Sonakshi Garg, Vicenç Torra
SEC2
2024 Explainability and Privacy-Preserving Data-Driven Models
Vicenç Torra
SECRYPT1
2024 Privacy in manifolds: Combining k-anonymity with differential privacy on Fréchet means
abstract
While anonymization techniques have improved greatly in allowing data to be used again, it is still really hard to get useful information from anonymized data without risking people’s privacy. Conventional approaches such as k-Anonymity and Differential Privacy have limitations in preserving data utility and privacy simultaneously, particularly in high-dimensional spaces with manifold structures. We address this challenge by focusing on anonymizing data existing within high-dimensional spaces possessing manifold structures. To tackle these issues, we propose and implement a hybrid anonymization scheme termed as the (β, k, b)-anonymization method that combines elements of both differential privacy and k-anonymity. This approach aims to produce high-quality anonymized data that closely resembles real data in terms of knowledge extraction while safeguarding privacy. The Fréchet mean, an operation applicable in metric spaces and meaningful in the manifold setting, serves as a key aspect of our approach. It provides insight into the geometry of data points within high-dimensional spaces. Our goal is to anonymize this Fréchet mean using our proposed approach and minimize the distance between the original and anonymized Fréchet mean to achieve data privacy without significant loss of information. Additionally, we introduce a novel Fréchet mean clustering model designed to enhance the clustering process for high-dimensional spaces. Through theoretical analysis and practical experiments, we demonstrate that our approach outperforms traditional privacy models both in terms of preserving data utility and privacy. This research contributes to advancing privacy-preserving techniques for complex and non-linear data structures, ensuring a balance between data utility and privacy protection.
Sonakshi Garg, Vicenç Torra
Comput. Secur.2
2024 Energy disaggregation risk resilience through microaggregation and discrete Fourier transform
abstract
Progress in the field of Non-Intrusive Load Monitoring (NILM) has been attributed to the rise in the application of artificial intelligence. Nevertheless, the ability of energy disaggregation algorithms to disaggregate different appliance signatures from aggregated smart grid data poses some privacy issues. This paper introduces a new notion of disclosure risk termed energy disaggregation risk. The performance of Sequence-to-Sequence (Seq2Seq) NILM deep learning algorithm along with three activation extraction methods are studied using two publicly available datasets. To understand the extent of disclosure, we study three inference attacks on aggregated data. The results show that Variance Sensitive Thresholding (VST) event detection method outperformed the other two methods in revealing households' lifestyles based on the signature of the appliances. To reduce energy disaggregation risk, we investigate the performance of two privacy-preserving mechanisms based on microaggregation and Discrete Fourier Transform (DFT). Empirically, for the first scenario of inference attack on UK-DALE, VST produces disaggregation risks of 99%, 100%, 89% and 99% for fridge, dish washer, microwave, and kettle respectively. For washing machine, Activation Time Extraction (ATE) method produces a disaggregation risk of 87%. We obtain similar results for other inference attack scenarios and the risk reduces using the two privacy-protection mechanisms.
Kayode S. Adewole, Vicenç Torra
Inf. Sci.2
2024 Computation of Choquet integrals: Analytical approach for continuous functions
abstract
In the continuous case, analytical computations of the Choquet integral are limited, despite being commonly used in various applications. One can either use the definition, which is computationally demanding and impractical, or apply already existing formulas restricted only to monotone nonnegative functions on a real interval starting at zero. This article aims to present more convenient computational formulas for continuous functions without imposing restrictions on their monotonicity given any real interval. First, a more general approach to monotone functions is provided for both positive and negative functions. Then, reordering techniques are introduced to compute the Choquet integral of an arbitrary continuous function, and with these, a monotone equivalent to every function can be constructed. This equivalent function preserves the final Choquet integral value, implying that only formulas for monotone functions are required. In addition to general fuzzy measures, the article assumes particular cases of distorted Lebesgue measures and distorted probabilities as the most commonly used fuzzy measures.
Zuzana Ontkovicová, Vicenç Torra
Inf. Sci.2
2024 $\Upsilon$-Values: Power Indices à La Orness for Nonadditive Measures
abstract
Fuzzy measures, also known as capacities, non-additive measures, and monotonic games, are increasingly used in all kind of applications. Fuzzy measures are set functions. So, for a given set$X$, we need to define$2^{|X|} - 2$parameters (excluding the measure on the empty set and on$X$itself). Because of that they are difficult to visualize, and indices and metrics have been defined. The Shapley value is an example. It permits us to determine weights of importance of each element in$X$. In this paper we introduce an alternative index. We call it$\Upsilon$-values. We provide an axiomatic characterization. These values are inspired on the Shapley values, and they are associated to set size, or position (order statistics) in a chain. Thus, also position when the measure is used in combination with a fuzzy integral. Andness and orness are measures that permit to evaluate the degree of simultaneity (conjunction) and substitutability (disjunction) of an aggregation function. We show the connection between our value and these concepts. In a way,$\Upsilon$-values define a power index à la orness.
Vicenç Torra
IEEE Trans. Fuzzy Syst.1
2023 Integrally Private Model Selection for Deep Neural Networks
Ayush K. Varshney, Vicenç Torra
DEXA (2)2
2023 Logic Aggregators and Their Implementations
Jozo J. Dujmovic, Vicenç Torra
MDAI2
2023 Differentially Private Graph Publishing Through Noise-Graph Addition
Julián Salas, Vladimiro González-Zelaya, Vicenç Torra, David Megías 0001
MDAI3
2023 Privacy Protection of Synthetic Smart Grid Data Simulated via Generative Adversarial Networks
abstract
The development in smart meter technology has made grid operations more efficient based on fine-grained electricity usage data generated at different levels of time granularity. Consequently, machine learning algorithms have benefited from these data to produce useful models for important grid operations. Although machine learning algorithms need historical data to improve predictive performance, these data are not readily available for public utilization due to privacy issues. The existing smart grid data simulation frameworks generate grid data with implicit privacy concerns since the data are simulated from a few real energy consumptions that are publicly available. This paper addresses two issues in smart grid. First, it assesses the level of privacy violation with the individual household appliances based on synthetic household aggregate loads consumption. Second, based on the findings, it proposes two privacy-preserving mechanisms to reduce this risk. Three inference attacks are simulated and the results obtained confirm the efficacy of the proposed privacy-preserving mechanisms.
Kayode S. Adewole, Vicenç Torra
SECRYPT2
2023 K-Anonymous Privacy Preserving Manifold Learning
abstract
In this modern world of digitalization, abundant amount of data is being generated. This often leads to data of high dimension, making data points far-away from each other. Such data may contain confidential information and must be protected from disclosure. Preserving privacy of this high-dimensional data is still a challenging problem. This paper aims to provide a privacy preserving model to anonymize high-dimensional data maintaining the manifold structure of the data. Manifold Learning hypothesize that real-world data lie on a low-dimensional manifold embedded in a higher-dimensional space. This paper proposes a novel approach that uses geodesic distance in manifold learning methods such as ISOMAP and LLE to preserve the manifold structure on low-dimensional embedding. Later on, anonymization of such sensitive data is achieved by M-MDAV, the manifold version of MDAV using geodesic distance. MDAV is a micro-aggregation privacy model. Finally, to evaluate the efficiency of the prop osed approach machine learning classification is performed on the anonymized lower-embedding. To emphasize the importance of geodesic-manifold learning, we compared our approach with a baseline method in which we try to anonymise high-dimensional data directly without reducing it onto a lower-dimensional space. We evaluate the proposed approach over natural and synthetic data such as tabular, image and textual data sets, and then empirically evaluate the performance of the proposed approach using different evaluation metrics viz. accuracy, precision, recall and K-Stress. We show that our proposed approach is providing accuracy up to 99% and thus, provides a novel contribution of analysing the effects of K-anonymity in manifold learning.
Sonakshi Garg, Vicenç Torra
SECRYPT2
2023 Δ SFL: (Decoupled Server Federated Learning) to Utilize DLG Attacks in Federated Learning by Decoupling the Server
abstract
Federated Learning or FL is the orchestration of centrally connected devices where a pre-trained machine learning model is sent to the devices and the devices train the machine learning model with their own data, individually. Though the data is not being stored in a central database the framework is still prone to data leakage or privacy breach. There are several different privacy attacks on FL such as, membership inference attack, gradient inversion attack, data poisoning attack, backdoor attack, deep learning from gradients attack (DLG). So far different technologies such as differential privacy, secure multi party computation, homomorphic encryption, k-anonymity etc. have been used to tackle the privacy breach. Nevertheless, there is very little exploration on the privacy by design approach and the analysis of the underlying network structure of the seemingly unrelated FL network. Here we are proposing the ΔDSFL framework, where the server is being decoupled into server and an an alyst. Also, in the learning process, ΔDSFL will learn the spatio information from the community detection, and then from DLG attack. Using the knowledge from both the algorithms, ΔDSFL will improve itself. We experimented on three different datasets (geolife trajectory, cora, citeseer) with satisfactory results.
Sudipta Paul 0008, Vicenç Torra
SECRYPT2
2023 On the definition of probabilistic metric spaces by means of fuzzy measures
abstract
Metric spaces are defined in terms of a space and a metric, or distance. Probabilistic metric spaces are a useful extension of metric spaces where the distance is a distribution instead of a number. In this way, we can take into account uncertainty. Then, the triangle inequality is replaced by a condition based on triangle functions on the distributions. In this paper we introduce F-spaces. This is a new type of probabilistic metric spaces which is based on fuzzy measures (also known as non-additive measures and capacities). We prove some properties that describe which families of fuzzy measures are compatible with which type of triangle functions. Then, we show how we can use Sugeno, Choquet integrals, and, in general, any other fuzzy integral as a tool for building these spaces. We show how these results can be used to compute distances between functions. We illustrate the example comparing three types of means when applied to a set of databases. The example uses Sugeno λ-measures to illustrate the theoretical results presented in the paper.
Yasuo Narukawa, Mariam Taha, Vicenç Torra
Fuzzy Sets Syst.3
2023 A systematic construction of non-i.i.d. data sets from a single data set: non-identically distributed data
abstract
Abstract Data-driven models strongly depend on data. Nevertheless, for research and academic purposes, public data sets are usually considered and analyzed. For example, most machine learning algorithms are applied and tested using the UCI Machine Learning repository. There is a current need for not i.i.d. data sets for distributed machine learning. Recall that i.i.d. random variables stand for independent and identically distributed (i.i.d.) random variables. An example of this need is federated learning. In federated learning, the typical scenario is to consider a set of agents each one with its own data set. Agents are typically heterogeneous and because of that, it is not appropriate to consider that the data of these agents follow the same distributions. In this paper we propose an approach to build non-identically distributed data sets from a single data set for machine learning classification, where we may suppose or not that all instances follow the same distribution. Each device will have only instances of a subset of the classes. The approach uses optimization to distribute the data set into a set of subsets, each one following a different distribution. Our goal is to define an approach for building subsets for training that is as systematic as the approaches used for cross-validation/k-fold validation.
Vicenç Torra
Knowl. Inf. Syst.1
2023 Scores for Hesitant Fuzzy Sets: Aggregation Functions and Generalized Integrals
abstract
There are several extensions of fuzzy sets. Hesitant fuzzy sets are one of them. They are defined in terms of a set of membership degrees. For a typical hesitant fuzzy set, this set of membership degrees has a finite number of values. One of the motivations to introduce score functions was to rank alternatives. In this case, as the membership degrees are a set, the comparison of membership values does not lead, in general, to a total order. Score functions can be seen as functions that transform the set of membership degrees into a single membership value. In this way, we can construct the total order. In this article, we propose a general framework to define score functions for hesitant fuzzy sets based on fuzzy integrals. This framework permits to see most relevant indices as particular cases. Moreover, previous approaches focused on typical hesitant fuzzy sets. Our approach is more general in the sense that we can process both typical hesitant fuzzy sets and the nontypical ones (with membership values that are not finite). We also frame the problem into a more general setting. That is, the problem of hesitant fuzzy set transformation. Score functions can be seen as functions that transform a hesitant fuzzy set into a standard fuzzy set. Similarly, we can consider its transformation to interval-valued fuzzy sets and type-2 fuzzy sets. Aggregation functions can also be used for the same purpose.
Yasuo Narukawa, Vicenç Torra
IEEE Trans. Fuzzy Syst.2
2022 Privacy Issues in Smart Grid Data: From Energy Disaggregation to Disclosure Risk
Kayode S. Adewole, Vicenç Torra
DEXA (1)2
2022 Designing Distributed Chi-Fuzzy Rule based Classification System
abstract
Fuzzy Rule based Classification Systems (FRBCSs) are an important area of fuzzy logic and fuzzy sets. FRBCSs provides interpretable models with good classification rate. In the presence of large number of instances with high dimension in training data, the classical Chi’s FRBCSs’ rulebase becomes huge causing the classical model to be computationally expensive which also leads to the degraded interpretability. In this work, we propose the ‘Distributed Chi Fuzzy Rule Based Classification Systems’ (DCHI-FRBCSs). The proposed approach distributes the n-dimensional data into n-nodes where each node forms its own 1-dimensional rules for the dimension. The output from each node is the rule weights which are then aggregated to form the final rule weights, the class with highest final rule weight is the final output of the model. Aggregators play an important role in combining the information from several sources. Andness-directed OWA selects desired level of andness between attributes for OWA aggregator. In this paper, we explore an extension of DCHI-FRBCSs called ‘Andess-directed Distributed Chi Fuzzy Rule Based Classification System’ (α-DCHI-FRBCS) which incorporates andness-directed OWA operator as the aggregator function. Results over UCI datasets verify the superiority of our methodology.
Ayush K. Varshney, Vicenç Torra
FUZZ-IEEE2
2022 Towards Integrally Private Clustering: Overlapping Clusters for High Privacy Guarantees
Vicenç Torra
PSD1
2022 (Max, ⊕)-transforms and genetic algorithms for fuzzy measure identification
abstract
Fuzzy measures generalize additive measures and probabilities. Their advantage with respect to additive ones is that they permit to model interactions between objects. Mesiar introduced in 1999 k-order Pan-additive fuzzy measures that generalize k-order additive and k-maxitive ones. They are related to the Möbius transform and related generalizations. In this paper we introduce some other transforms that we call (Max,+) and (Max,⊕) that permit to represent fuzzy measures in a convenient way when we use genetic algorithms in fuzzy measure identification problems. We illustrate its use identifying a measure for a subjective evaluation problem using a Choquet integral and a Sugeno integral.
Vicenç Torra
Fuzzy Sets Syst.1
2022 Preface: Fuzziness and Uncertainty in Data Mining
abstract
International Journal of Uncertainty, Fuzziness and Knowledge-Based SystemsVol. 30, No. Supp02, pp. v (2022) No AccessPreface: Fuzziness and Uncertainty in Data MiningVicenç Torra and Yasuo NarukawaVicenç TorraDepartment of Computing Sciences, Umeå University, Sweden and Yasuo NarukawaDepartment of Management Science, Tamagawa University, 6-1-1 Tamagawagakuen, Machida, Tokyo, 194-8610, Japanhttps://doi.org/10.1142/S0218488522020044Cited by:0 Next This article is part of the issue: Special Issue on Fuzziness and Uncertainty in Data Mining AboutSectionsPDF/EPUB ToolsAdd to favoritesDownload CitationsTrack CitationsRecommend to Library ShareShare onFacebookTwitterLinked InRedditEmail Remember to check out the Most Cited Articles! Check out our titles on Fuzzy Logic & Z-Numbers With a wide range of areas, you're bound to find something you like. FiguresReferencesRelatedDetails Recommended Vol. 30, No. Supp02 Metrics History PDF download
Vicenç Torra, Yasuo Narukawa
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2021 A Survey on Tree Aggregation
abstract
The research dedicated to the aggregation of classification trees and general trees (hierarchical structure of objects) has made enormous progress in the past decade. The problem statement for aggregation of classification trees or general trees is as follows: Given k classification or general trees for a set of objects, we aim to build a consensus tree (classification or general). That is, a representative tree for the given trees. In this paper, we explore different perspectives for the motivation to construct a single tree from multiple trees given by researchers. The survey presents the approaches for the aggregation of both the classification trees as well as general trees. We bifurcate our study of the aggregation approaches into two categories: Selecting a single tree from multiple trees and merging trees. We will discuss these categories and the aggregation approaches under these categories in the paper comprehensively. We also discuss the privacy aspects of tree aggregation approaches and the possible directions for new research like using the technique of aggregating decision trees in the field of Federated Learning, which is a booming topic.
Saloni Kwatra, Vicenç Torra
FUZZ-IEEE2
2021 An score index for hesitant fuzzy sets based on the Choquet integral
abstract
In a previous paper we introduced hesitant fuzzy sets. A typical hesitant fuzzy set (HFS) assigns a finite set of membership values to each element. Because of that, in general, the comparison of the membership of different elements to the same set does not lead to a total order. In this paper we study the definition of score functions. Score functions were introduced for HFS. They are functions that summarize the membership values of a given element to the fuzzy set and allows us to build this total order for hesitant fuzzy sets. We propose a new index based on the Choquet integral and we compare this approach with some existing indices in the literature.
Yasuo Narukawa, Vicenç Torra
FUZZ-IEEE2
2021 Systematic Evaluation of Probabilistic k-Anonymity for Privacy Preserving Micro-data Publishing and Analysis
abstract
In the light of stringent privacy laws, data anonymization not only supports privacy preserving data publication (PPDP) but also improves the flexibility of micro-data analysis. Machine learning (ML) is widely used for personal data analysis in the present day thus, it is paramount to understand how to effectively use data anonymization in the ML context. In this work, we introduce an anonymization framework based on the notion of “probabilistic k-anonymity” that can be applied with respect to mixed datasets while addressing the challenges brought forward by the existing syntactic privacy models in the context of ML. Through systematic empirical evaluation, we show that the proposed approach can effectively limit the disclosure risk in micro-data publishing while maintaining a high utility for the ML models induced from the anonymized data.
Navoda Senavirathne, Vicenç Torra
SECRYPT2
2021 Andness directedness for operators of the OWA and WOWA families
abstract
Andness directed aggregation is about selecting aggregators from a desired andness level. In this paper we consider operators of the OWA and WOWA families: aggregation functions that permit us to represent some degree of compensation of the input values. In addition to compensation, WOWA permits us to represent importance (weights) of the input values. Selection of appropriate parameters given an andness level will be based on families of fuzzy quantifiers.
Vicenç Torra
Fuzzy Sets Syst.1
2021 Properties and comparison of andness-characterized aggregators
abstract
Logic aggregators are all aggregators characterized by andness/orness. In this paper we present necessary properties of logic aggregators and use them to compare major implementations of andness-characterized aggregators: means, t-norms/conorms, ordered weighted average family, fuzzy integrals, and graded conjunction/disjunction. Our goal is to provide methodology for justifiable selection of the most suitable aggregator for various information fusion problems. When aggregating arguments from [0, 1], such arguments are regularly interpreted as degrees of truth or degrees of fuzzy membership. In such cases we deal with logic aggregators that must have appropriate logic properties. Our analysis identifies 10 necessary properties that are critical for decision support applications and must be satisfied by logic aggregators. Then, we evaluate and compare five families of logic aggregators that offer different levels of support to desired logic properties.
Jozo J. Dujmovic, Vicenç Torra
Int. J. Intell. Syst.2
2020 Explaining Recurrent Machine Learning Models: Integral Privacy Revisited
abstract
Abstract We have recently introduced a privacy model for statistical and machine learning models called integral privacy. A model extracted from a database or, in general, the output of a function satisfies integral privacy when the number of generators of this model is sufficiently large and diverse. In this paper we show how the maximal c-consensus meets problem can be used to study the databases that generate an integrally private solution. We also introduce a definition of integral privacy based on minimal sets in terms of this maximal c-consensus meets problem.
Vicenç Torra, Guillermo Navarro-Arribas, Edgar Galván López
PSD1
2020 On the Role of Data Anonymization in Machine Learning Privacy
abstract
Data anonymization irrecoverably transforms the raw data into a protected version by eliminating direct identifiers and removing sufficient details from indirect identifiers in order to minimize the risk of re-identification when there is a requirement for data publishing. Nevertheless, data protection laws (i.e., GDPR) do not consider anonymized data as personal data thus allowing them to be freely used, analysed, shared and monetized without a compliance risk. Motivated by the above advantages, it is plausible that the data controllers anonymize the data before releasing them for any data analysis tasks such as machine learning (ML); which is applied in a wide variety of domains where personal data are used. Moreover, in recent research, it has shown that ML models are vulnerable to privacy attacks as they retain sensitive information from the training data. Taking all of these facts into consideration, in this work we explore the interplay between data anonymization and ML with the ultimate aim of clarifying whether data anonymization is sufficient to achieve privacy for ML under different adversarial scenarios. We also discuss the challenges and opportunities of integrating these two domains. As per our findings, it is conspicuous that in order to substantially minimize the privacy risks in ML, existing data anonymization techniques have to be applied with high privacy levels that cause a deterioration in model utility.
Navoda Senavirathne, Vicenç Torra
TrustCom2
2020 On the f-divergence for discrete non-additive measures
Vicenç Torra, Yasuo Narukawa, Michio Sugeno
Inf. Sci.1
2020 Swapping trajectories with a sufficient sanitizer
abstract
Real-time mobility data is useful for several applications such as planning transports in metropolitan areas or localizing services in towns. However, if such data is collected without any privacy protection it may reveal sensible locations and pose safety risks to an individual associated to it. Thus, mobility data must be anonymized preferably at the time of collection. In this paper, we consider the SwapMob algorithm that mitigates privacy risks by swapping partial trajectories. We formalize the concept of sufficient sanitizer and show that the SwapMob algorithm is a sufficient sanitizer for various statistical decision problems. That is, it preserves the aggregate information of the spatial database in the form of sufficient statistics and also provides privacy to the individuals. This may be used for personalized assistants taking advantage of users’ locations, so they can ensure user privacy while providing accurate response to the user requirements. We measure the privacy provided by SwapMob as the Adversary Information Gain, which measures the capability of an adversary to leverage his knowledge of exact data points to infer a larger segment of the sanitized trajectory. We test the utility of the data obtained after applying SwapMob sanitization in terms of Origin-Destination matrices, a fundamental tool in transportation modelling.
Julián Salas, David Megías 0001, Vicenç Torra, Marina Toger, Joel Dahne, Raazesh Sainudiin
Pattern Recognit. Lett.3
2019 Managing the doubt in fuzzy clustering by means of interval-valued fuzzy sets
abstract
In this work we study how the outliers can distort a partitional clustering process. We present a new algorithm to avoid this distortion. It is based on the minimization of a new objective functions, which is an extension of the one of the Fuzzy Clusters Means algorithm. The main novelty is the use of interval values to calculate the membership degrees of each datum to each cluster. We show the performance of our proposal over different datasets and we present its advantages in image segmentation.
Aranzazu Jurio, Humberto Bustince, Vicenç Torra
FUZZ-IEEE3
2019 Derivative for Discrete Choquet Integrals
Yasuo Narukawa, Vicenç Torra
MDAI2
2019 Towards an Adaptive Defuzzification: Using Numerical Choquet Integral
Vicenç Torra, Joaquín García 0001
MDAI1
2019 Integrally private model selection for decision trees
abstract
Privacy attacks targeting machine learning models are evolving. One of the primary goals of such attacks is to infer information about the training data used to construct the models. “Integral Privacy” focuses on machine learning and statistical models which explain how we can utilize intruder’s uncertainty to provide a privacy guarantee against model comparison attacks. Through experimental results, we show how the distribution of models can be used to achieve integral privacy. Here, we observe two categories of machine learning models based on their frequency of occurrence in the model space. Then we explain the privacy implications of selecting each of them based on a new attack model and empirical results. Also, we provide recommendations for private model selection based on the accuracy and stability of the models along with the diversity of training data that can be used to generate the models.
Navoda Senavirathne, Vicenç Torra
Comput. Secur.2
2019 A note on some algebraic properties of discrete Sugeno integrals
Radomír Halas, Radko Mesiar, Jozef Pócs, Vicenç Torra
Fuzzy Sets Syst.4
2019 A Hierarchically ⊥-Decomposable Fuzzy Measure-Based Approach for Fuzzy Rules Aggregation
abstract
A Fuzzy Decision Tree is a classification method consisting of a set of rules defined on fuzzy variables. The final class assignment is done according to the output of all the rules of the tree. Generally, the maximum operator is used to aggregate the results of the rules. However, some approaches based on more complex aggregation operators have appeared recently. In this work we propose to use Sugeno and Choquet integrals together with a Hierarchically ⊥-Decomposable Fuzzy Measure (HDFM) to aggregate the rules' values. The HDFM exploits the hierarchical structure of the fuzzy decision tree and takes into account the confidence value of the output together with the classification ambiguity of the rules. The HDFM is built using Sugeno-Weber t-conorms.We validate this approach on several classification problems and make a comparison of the performance with the state of art aggregation operators. Finally, a case study with a real dataset of diabetic patients is analyzed to predict the risk of suffering from diabetic retinopathy.
Emran Saleh, Aïda Valls, Antonio Moreno, Pedro Romero-Aroca, Humberto Bustince, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.6
2019 Foreword: Tools for Decision Making under Vagueness
Vicenç Torra, Yasuo Narukawa
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2019 On network analysis using non-additive integrals: extending the game-theoretic network centrality
abstract
There are large amounts of information that can be represented in terms of graphs. This includes social networks and internet. We can represent people and their interactions by means of graphs. Similarly, we can represent web pages (and sites) as well as links between pages by means of graphs. In order to study the properties of graphs, several indices have been defined. They include degree centrality, betweenness, and closeness. In this paper, we propose the use of Choquet and Sugeno integrals with respect to non-additive measures for network analysis. This is a natural extension of the use of game theory for network analysis. Recall that monotonic games in game theory are non-additive measures. We discuss the expected force, a centrality measure, in the light of non-additive integral network analysis. We also show that some results by Godo et al. can be used to compute network indices when the information associated with a graph is qualitative.
Vicenç Torra, Yasuo Narukawa
Soft Comput.1
2018 Learning Fuzzy Measures for Aggregation in Fuzzy Rule-Based Models
Emran Saleh, Aïda Valls, Antonio Moreno, Pedro Romero-Aroca, Vicenç Torra, Humberto Bustince
MDAI5
2018 SwapMob: Swapping Trajectories for Mobility Anonymization
abstract
Abstract Mobility data mining can improve decision making, from planning transports in metropolitan areas to localizing services in towns. However, unrestricted access to such data may reveal sensible locations and pose safety risks if the data is associated to a specific moving individual. This is one of the many reasons to consider trajectory anonymization. Some anonymization methods rely on grouping individual registers on a database and publishing summaries in such a way that individual information is protected inside the group. Other approaches consist of adding noise, such as differential privacy, in a way that the presence of an individual cannot be inferred from the data. In this paper, we present a perturbative anonymization method based on swapping segments for trajectory data (SwapMob). It preserves the aggregate information of the spatial database and at the same time, provides anonymity to the individuals. We have performed tests on a set of GPS trajectories of 10,357 taxis during the period of Feb. 2 to Feb. 8, 2008, within Beijing. We show that home addresses and POIs of specific individuals cannot be inferred after anonymizing them with SwapMob, and remark that the aggregate mobility data is preserved without changes, such as the average length of trajectories or the number of cars and their directions on any given zone at a specific time.
Julián Salas, David Megías 0001, Vicenç Torra
PSD3
2018 Approximating Robust Linear Regression With An Integral Privacy Guarantee
abstract
Most of the privacy-preserving techniques suffer from an inevitable utility loss due to different perturbations carried out on the input data or the models in order to gain privacy. When it comes to machine learning (ML) based prediction models, accuracy is the key criterion for model selection. Thus, an accuracy loss due to privacy implementations is undesirable.The motivation of this work, is to implement the privacy model “integral privacy” and to evaluate its eligibility as a technique for machine learning model selection while preserving model utility. In this paper, a linear regression approximation method is implemented based on integral privacy which ensures high accuracy and robustness while maintaining a degree of privacy for ML models. The proposed method uses a re-sampling based estimator to construct linear regression model which is coupled with a rounding based data discretization method to support integral privacy principles. The implementation is evaluated in comparison with differential privacy in terms of privacy, accuracy and robustness of the output ML models. In comparison, integral privacy based solution provides a better solution with respect to the above criteria.
Navoda Senavirathne, Vicenç Torra
PST2
2018 Synthetic generation of spatial graphs
abstract
Graphs can be used to model many different types of interaction networks, for example, online social networks or animal transport networks. Several algorithms have thus been introduced to build graphs according to some predefined conditions. In this paper, we present an algorithm that generates spatial graphs with a given degree sequence. In spatial graphs, nodes are located in a space equiped with a metric. Our goal is to define a graph in such a way that the nodes and edges are positioned according to an underlying metric. More particularly, we have constructed a greedy algorithm that generates nodes proportional to an underlying probability distribution from the spatial structure, and then generates edges inversely proportional to the Euclidean distance between nodes. The algorithm first generates a graph that can be a multigraph, and then corrects multiedges. Our motivation is in data privacy for social networks, where a key problem is the ability to build synthetic graphs. These graphs need to satisfy a set of required properties (e.g., the degrees of the nodes) but also be realistic, and thus, nodes (individuals) should be located according to a spatial structure and connections should be added taking into account nearness.
Vicenç Torra, Annie Jonsson, Guillermo Navarro-Arribas, Julián Salas
Int. J. Intell. Syst.1
2017 Provenance and Privacy
Vicenç Torra, Guillermo Navarro-Arribas, David Sanchez-Charles, Victor Muntés-Mulero
MDAI1
2017 Entropy for non-additive measures in continuous domains
Vicenç Torra
Fuzzy Sets Syst.1
2017 k-Degree anonymity and edge selection: improving data utility in large networks
Jordi Casas-Roma, Jordi Herrera-Joancomartí, Vicenç Torra
Knowl. Inf. Syst.3
2016 Integral Privacy
Vicenç Torra, Guillermo Navarro-Arribas
CANS1
2016 Fuzzy, I-fuzzy, and H-fuzzy partitions to describe clusters
abstract
In this paper we discuss how three types of fuzzy partitions can be used to describe the results of three types of cluster structures. Standard fuzzy partitions are suitable for centroid based clusters, and I-fuzzy partitions for clusters represented by segments or lines (e.g., c-varieties). In this paper, we introduce hesitant fuzzy partitions. They are suitable for clusters defined by sets of centroids. Because of that, we show that they are useful for hierarchical clustering. We also establish the relationship between hesitant fuzzy partitions and I-fuzzy partitions.
Vicenç Torra, Laya Aliahmadipour, Anders Dahlbom
FUZZ-IEEE1
2016 Partial Domain Theories for Privacy
Eva Armengol, Vicenç Torra
MDAI2
2016 Improving the characterization of P-stability for applications in network privacy
Julián Salas, Vicenç Torra
Discret. Appl. Math.2
2016 On the f-divergence for non-additive measures
Vicenç Torra, Yasuo Narukawa, Michio Sugeno
Fuzzy Sets Syst.1
2016 On a Relationship Between Fuzzy Measures and AIFS
abstract
The literature discusses several extensions of fuzzy sets. AIFS, IVFS, HFS, type-2 fuzzy sets are some of them. Interval valued fuzzy sets is one of the extensions where the membership is not a single value but an interval. Atanassov Intuitionistic fuzzy sets, for short AIFS, are defined in terms of two values for each element: membership and non-membership. In this paper we discuss AIFS and their relationship with fuzzy measures. The discussion permits us to define counter AIFS (cIFS) and discretionary AIFS (dIFS). They are extensions of fuzzy sets that are based on fuzzy measures.
Vicenç Torra, Yasuo Narukawa, Ronald R. Yager
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2015 Generalization-Based k-Anonymization
Eva Armengol, Vicenç Torra
MDAI2
2015 Spherical microaggregation: Anonymizing sparse vector spaces
Daniel Abril, Guillermo Navarro-Arribas, Vicenç Torra
Comput. Secur.3
2015 Graphic sequences, distances and k-degree anonymity
Julián Salas, Vicenç Torra
Discret. Appl. Math.2
2015 Anonymizing graphs: measuring quality for clustering
Jordi Casas-Roma, Jordi Herrera-Joancomartí, Vicenç Torra
Knowl. Inf. Syst.3
2014 Comparing fuzzy measures through their Möbius transform
Vicenç Torra, Yasuo Narukawa, Daniel Abril
FUSION1
2014 Choquet Integral on Multisets
Yasuo Narukawa, Vicenç Torra
IPMU (1)2
2014 Rank Swapping for Stream Data
Guillermo Navarro-Arribas, Vicenç Torra
MDAI2
2014 JPEG-Based Microdata Protection
Javier Jiménez, Guillermo Navarro-Arribas, Vicenç Torra
Privacy in Statistical Databases3
2014 Hesitant Fuzzy Sets: An Emerging Tool in Decision Making
Francisco Herrera, Luis Martínez-López 0001, Vicenç Torra, Zeshui Xu
Int. J. Intell. Syst.3
2014 Hesitant Fuzzy Sets: State of the Art and Future Directions
abstract
The necessity of dealing with uncertainty in real world problems has been a long-term research challenge that has originated different methodologies and theories. Fuzzy sets along with their extensions, such as type-2 fuzzy sets, interval-valued fuzzy sets, and Atanassov's intuitionistic fuzzy sets, have provided a wide range of tools that are able to deal with uncertainty in different types of problems. Recently, a new extension of fuzzy sets so-called hesitant fuzzy sets has been introduced to deal with hesitant situations, which were not well managed by the previous tools. Hesitant fuzzy sets have attracted very quickly the attention of many researchers that have proposed diverse extensions, several types of operators to compute with such types of information, and eventually some applications have been developed. Because of such a growth, this paper presents an overview on hesitant fuzzy sets with the aim of providing a clear perspective on the different concepts, tools and trends related to this extension of fuzzy sets.
Rosa M. Rodríguez 0001, Luis Martínez-López 0001, Vicenç Torra, Zeshui Xu, Francisco Herrera
Int. J. Intell. Syst.3
2014 Aggregation functions for typical hesitant fuzzy elements and the action of automorphisms
Benjamín R. C. Bedregal, Renata H. S. Reiser, Humberto Bustince, Carlos Lopez-Molina, Vicenç Torra
Inf. Sci.5
2014 An evolutionary algorithm to enhance multivariate Post-Randomization Method (PRAM) protections
Jordi Marés, Vicenç Torra
Inf. Sci.2
2013 An algorithm for k-degree anonymity on large networks
abstract
In this paper, we consider the problem of anonymization on large networks. There are some anonymization methods for networks, but most of them can not be applied on large networks because of their complexity. We present an algorithm for k-degree anonymity on large networks. Given a network G, we construct a k-degree anonymous network, G, by the minimum number of edge modifications. We devise a simple and efficient algorithm for solving this problem on large networks. Our algorithm uses univariate micro-aggregation to anonymize the degree sequence, and then it modifies the graph structure to meet the k-degree anonymous sequence. We apply our algorithm to a different large real datasets and demonstrate their efficiency and practical utility.
Jordi Casas-Roma, Jordi Herrera-Joancomartí, Vicenç Torra
ASONAM3
2013 Analyzing the Impact of Edge Modifications on Networks
Jordi Casas-Roma, Jordi Herrera-Joancomartí, Vicenç Torra
MDAI3
2013 A self-adaptive classification for the dissociating privacy agent
abstract
This paper describes an extension of the Dissociating Privacy Agent (DisPA), which is a Privacy Enhancement Technology (PET) for web search. It is implemented as an add-on for Firefox that acts like a proxy between the user and the search engine. A fundamental part of DisPA is the classification of queries into a set of categories. Particularly, the taxonomy of the Open Directory Project (ODP) is used for this purpose. In this paper we briefly recall the internal operations of the agent, discuss the drawbacks of the current model and propose an improvement to overcome them.
Marc Juarez, Vicenç Torra
PST2
2013 Toward a Privacy Agent for Information Retrieval
abstract
In this paper, we tackle the private information retrieval (PIR) problem associated with the use of Internet search engines. We address the desire for a user to retrieve information from the Web without the search provider learning about it. Traditional PIR protocols present two main shortcomings for their application: (i) They assume cooperation by the database, which is not affordable for a real-world search engine like Google and (ii) their computational complexity is linear in the size of the database, which is unfeasible in the case of the Web. More recent approaches relax PIR conditions to overcome these limitations and present some level of privacy. Mostly, they aim to distort server logs regardless of the loss of information that is involved. Server logs are used by search engines for profiling and, thereby, provide personalized results. This becomes a user's need given the growth of the Web and can also be used for targeted advertising. This study focuses on a noncooperative agent for private search that considers profiling as valuable data used for both sides of the search process. It is based on the assumption that the user's identity is formed by the union of various areas of interests or facets. Managing the HTTP connections properly, submitted queries are mapped to different server logs according to these facets. The rationale is that these logs cannot be used for tracing the user while they are still helpful for profiling. We present a personalized query classification approach based on the user's browsing history and to provide empirical results; we developed an attacking algorithm against the agent that shows that the disclosure risk is reduced.
Marc Juarez, Vicenç Torra
Int. J. Intell. Syst.2
2013 Technologies for Decision Making and AI Applications
Vicenç Torra, Yasuo Narukawa, Jianping Yin
Int. J. Intell. Syst.1
2013 On the protection of social networks user's information
Jordi Marés, Vicenç Torra
Knowl. Based Syst.2
2012 On the relationship between clustering and coding theory
abstract
In this paper we discuss the relations between clustering and error correcting codes. We show that clustering can be used for constructing error correcting codes. We review the previous works found in the literature about this issue, and propose a modification of a previous work that can be used for code construction from a set of proposed codewords.
Klara Stokes, Vicenç Torra
FUZZ-IEEE2
2012 Comparing Random-Based and k-Anonymity-Based Algorithms for Graph Anonymization
Jordi Casas-Roma, Jordi Herrera-Joancomartí, Vicenç Torra
MDAI3
2012 Heuristic Supervised Approach for Record Linkage
Javier Murillo, Daniel Abril, Vicenç Torra
MDAI3
2012 Clustering-Based Categorical Data Protection
Jordi Marés, Vicenç Torra
Privacy in Statistical Databases2
2012 Guest Editors' Introduction
Klara Stokes, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2012 Multiple Releases of k-Anonymous Data Sets and k-Anonymous Relational Databases
abstract
In data privacy, the evaluation of the disclosure risk has to take into account the fact that several releases of the same or similar information about a population are common. In this paper we discuss this issue within the scope of k-anonymity. We also show how this issue is related to the publication of privacy protected databases that consist of linked tables. We present algorithms for the implementation of k-anonymity for this type of data.
Klara Stokes, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2012 A Formalization of Record Linkage and its Application to Data Protection
abstract
Re-identification and record linkage are tools used to measure disclosure risk in data privacy. Given two data files, record linkage establishes links between those records that correspond to the same individual. These links are often expressed in terms of probability distributions. This paper presents a review of a formalization of re-identification in terms of compatible belief functions. This formalization makes it possible to define the set of methods for re-identification that are relevant for the estimation of disclosure risk in privacy protection. Any re-identification method that does not fulfill the criteria for being in this set, may be discarded in a theoretical disclosure risk analysis. The focus in this paper is on providing examples of how this formalization can be applied in a few different scenarios in data privacy
Vicenç Torra, Klara Stokes
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2012 User k-anonymity for privacy preserving data mining of query logs
Guillermo Navarro-Arribas, Vicenç Torra, Arnau Erola, Jordi Castellà-Roca
Inf. Process. Manag.2
2012 On a comparison between Mahalanobis distance and Choquet integral: The Choquet-Mahalanobis operator
Vicenç Torra, Yasuo Narukawa
Inf. Sci.1
2012 Reidentification and k-anonymity: a model for disclosure risk in graphs
Klara Stokes, Vicenç Torra
Soft Comput.2
2012 An Extension of Fuzzy Measures to Multisets and Its Relation to Distorted Probabilities
abstract
Fuzzy measures are monotonic set functions on a reference set; they generalize probabilities replacing the additivity condition by monotonicity. The typical application of these measures is with fuzzy integrals. Fuzzy integrals integrate a function with respect to a fuzzy measure, and they can be used to aggregate information from a set of sources (opinions from experts or criteria in a multicriteria decision-making problem). In this context, background knowledge on the sources is represented by means of the fuzzy measures. For example, interactions between criteria are represented by means of nonadditive measures. In this paper, we introduce fuzzy measures on multisets. We propose a general definition, and we then introduce a family of fuzzy measures for multisets which we show to be equivalent to distorted probabilities when the multisets are restricted to proper sets.
Vicenç Torra, Klara Stokes, Yasuo Narukawa
IEEE Trans. Fuzzy Syst.1
2011 Trajectory anonymization from a time series perspective
abstract
In a world of constant technological evolution, the expansion of location tracking technologies is a fact. Most of us may have a device which can tell us our location with more or less accuracy. Most of those devices collect our location information and, usually, the trajectories we make. Publishing this information is essential for companies to improve their marketing strategies, for the traffic department in order to have some control about the traffic, and for many other entities. With the publishment of this information an important problem arises, the privacy. On the literature, many approaches to trajectory protection are provided. We contribute to the scene providing a framework with a protection method and measures for both the computation of the perturbation added to the data, and the risk of linking a protected trajectory to its original.
Sergi Martinez-Bea, Vicenç Torra
FUZZ-IEEE2
2011 On some clustering approaches for graphs
abstract
In this paper we discuss some tools for graph perturbation with applications to data privacy. We present and analyse two different approaches. One is based on matrix decomposition and the other on graph partitioning. We discuss these methods and show that they belong to two traditions in data protection: noise addition/microaggregation and k-anonymity.
Klara Stokes, Vicenç Torra
FUZZ-IEEE2
2011 On the Declassification of Confidential Documents
Daniel Abril, Guillermo Navarro-Arribas, Vicenç Torra
MDAI3
2011 Fuzzy Measures and Comonotonicity on Multisets
Yasuo Narukawa, Klara Stokes, Vicenç Torra
MDAI3
2011 A Comparison of Two Different Types of Online Social Network from a Data Privacy Perspective
David F. Nettleton, Diego Sáez-Trumper, Vicenç Torra
MDAI3
2011 On distorted probabilities and m-separable fuzzy measures
Yasuo Narukawa, Vicenç Torra
Int. J. Approx. Reason.2
2011 Computationally intensive parameter selection for clustering algorithms: The case of fuzzy c-means with tolerance
abstract
Parameter selection is a well-known problem in the fuzzy clustering community. In this paper, we propose to tackle this problem using a computationally intensive approach. We apply this approach to a new method for clustering recently introduced in the literature. It is the fuzzy c-means with tolerance. This method permits data to include some error, and this is modeled by moving data in a particular direction within a particular range when clusters are defined. The proper application of this approach needs the correct definition of the parameter κ. A value that might be different for each record and corresponds to the maximum shift allowed to the data. In this paper, we review this method and we study the definition of this parameter κ when the same value of κ is used for all data elements. Our approach is based on the analysis of sets of data with increasing noise and an exhaustive analysis of the behavior of the algorithm with different values of κ. The analysis is motivated in privacy preserving data mining. The same approach can be used for parameter selection in other clustering algorithms. © 2010 Wiley Periodicals, Inc.
Vicenç Torra, Yasunori Endo, Sadaaki Miyamoto
Int. J. Intell. Syst.1
2011 Data Privacy for Simply Anonymized Social Network Logs Represented as Graphs - Considerations for Graph alteration Operations
abstract
In this paper we review the state of the art on graph privacy with special emphasis on applications to online social networks, and we consider some novel aspects which have not been greatly covered in the specialized literature on graph privacy. The following key considerations are covered: (i) choice of different operators to modify the graph; (ii) information loss based on the cost of graph operations in terms of statistical characteristics (degree, clustering coefficient and path length) in the original graph; (iii) computational cost of the operations; (iv) in the case of the aggregation of two nodes, the choice of similar adjacent nodes rather than isomorphic topologies, in order to maintain the overall structure of the graph; (v) a statistically knowledgeable attacker who is able to search for regions of a simply anonymized graph based on statistical characteristics and map those onto a given node and its immediate neighborhood.
David F. Nettleton, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2011 Guest Editors' Introduction
abstract
Work partially funded by Spanish MEC, projects ARES – CONSOLIDER INGENIO 2010 CSD2007-00004 – and eAEGIS – TSI2007-65406-C03-02
Vicenç Torra, Yasuo Narukawa, Marc Daumas
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2011 Flexible secure inter-domain interoperability through attribute conversion
Carles Martínez-García, Guillermo Navarro-Arribas, Simon N. Foley, Vicenç Torra, Joan Borrell
Inf. Sci.4
2011 An evolutionary approach to enhance data privacy
Javier Jiménez, Jordi Marés, Vicenç Torra
Soft Comput.3
2011 A definition for I-fuzzy partitions
Vicenç Torra, Sadaaki Miyamoto
Soft Comput.1
2011 Erratum to: A definition for I-fuzzy partitions
Vicenç Torra, Sadaaki Miyamoto
Soft Comput.1
2010 Evaluation of information loss for privacy preserving data mining through comparison of fuzzy partitions
abstract
In this paper, we focus on the problem of preserving the data confidentiality when sharing the data for clustering. This problem poses new challenges for novel uses of privacy preserving data mining (PPDM) techniques. Specifically, this paper considers the synthetic data generation as a way to preserve the data privacy. One of the state of the art synthetic data generators is the IPSO family of methods. It has been stated that the use of IPSO to generate synthetic data is appropriate when the user plans to apply clustering to the data. Moreover, this paper aims to associate the same property to the FCRM synthetic data generator, and at the same time, to assess the relationship between the information loss produced when generating synthetic data with FCRM and the clustering similarity between the original and synthetic data.
Isaac Cano, Susana Ladra, Vicenç Torra
FUZZ-IEEE3
2010 Possibilistic reasoning for trust-based access control enforcement in social networks
abstract
Web 2.0 services allow large amount of users to interact and share resources with each others in a very easy and timely way. However, as such services are very recent, there is still a need for developing new tools that considers their own characteristics and particularities. This is the case of the resource access control methods. In this paper we present the Max-Min method, a new trust-based algorithm for computing user reputation in social/peer-to-peer networks in an easy, fast and comprehensible manner. For these reasons, it can be used for resource access control enforcement. The novelty of Max-Min is that it is specifically designed for Web 2.0 applications. Therefore, it considers all the particularities of such scenarios. Additionally, we include in this work a set of experiments showing the viability of Max-Min model using a real trust dataset downloaded from Epinions, a Web 2.0 website.
Jordi Nin, Vicenç Torra
FUZZ-IEEE2
2010 Towards Semantic Microaggregation of Categorical Data for Confidential Documents
Daniel Abril, Guillermo Navarro-Arribas, Vicenç Torra
MDAI3
2010 A Bibliometric Index Based on Collaboration Distances
Maria Bras-Amorós, Josep Domingo-Ferrer, Vicenç Torra
MDAI3
2010 Using Classification Methods to Evaluate Attribute Disclosure Risk
Jordi Nin, Javier Herranz, Vicenç Torra
MDAI3
2010 Semantic Microaggregation for the Anonymization of Query Logs
Arnau Erola, Jordi Castellà-Roca, Guillermo Navarro-Arribas, Vicenç Torra
Privacy in Statistical Databases4
2010 PRAM Optimization Using an Evolutionary Algorithm
Jordi Marés, Vicenç Torra
Privacy in Statistical Databases2
2010 Classifying data from protected statistical datasets
Javier Herranz, Stan Matwin, Jordi Nin, Vicenç Torra
Comput. Secur.4
2010 Hesitant fuzzy sets
abstract
Several extensions and generalizations of fuzzy sets have been introduced in the literature, for example, Atanassov's intuitionistic fuzzy sets, type 2 fuzzy sets, and fuzzy multisets. In this paper, we propose hesitant fuzzy sets. Although from a formal point of view, they can be seen as fuzzy multisets, we will show that their interpretation differs from the two existing approaches for fuzzy multisets. Because of this, together with their definition, we also introduce some basic operations. In addition, we also study their relationship with intuitionistic fuzzy sets. We prove that the envelope of the hesitant fuzzy sets is an intuitionistic fuzzy set. We prove also that the operations we propose are consistent with the ones of intuitionistic fuzzy sets when applied to the envelope of the hesitant fuzzy sets. © 2010 Wiley Periodicals, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
2010 Information Loss for Synthetic Data through Fuzzy Clustering
abstract
Synthetic data generators are one of the methods used in privacy preserving data mining for ensuring the privacy of the individuals when their data are published. Synthetic data generators construct artificial data from some models obtained from the original data. Such models are mainly based on statistics and, typically, do not take into account other aspects of interest in artificial intelligence. In this paper we study whether one family of such synthetic data generators (the IPSO family) preserves the properties of the data that are of interest when users plan to apply clustering techniques. In particular, we study the effect of such synthetic data generators on fuzzy clustering. That is, we study the information loss data suffer when the original data are replaced by the synthetic ones.
Susana Ladra, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2010 Container loading for nonorthogonal objects: an approximation using local search and simulated annealing
Vicenç Torra, Isaac Cano, Sadaaki Miyamoto, Yasunori Endo
Soft Comput.1
2010 Soft Computing in decision modeling
Vicenç Torra, Yasuo Narukawa
Soft Comput.1
2010 Some relationships between Losonczi's based OWA generalizations and the Choquet-Stieltjes integral
Vicenç Torra, Yasuo Narukawa
Soft Comput.1
2009 Utility and Risk of JPEG-Based Continuous Microdata Protection Methods
abstract
Releasing data containing sensitive information has an implicit risk that confidential information about individuals became revealed. Perturbative masking methods propose the distortion of the original data sets before publication, in order to obtain a tradeoff between data utility (low information loss) and protection against disclosure (low disclosure risk). In this paper, we empirically evaluate a particular collection of perturbative masking methods for continuous microdata based on the image compression standards JPEG and JPEG 2000.
Javier Jiménez, Vicenç Torra
ARES2
2009 Rank Swapping for Partial Orders and Continuous Variables
abstract
Rank swapping, which was first defined for ordinal attributes, is currently applied also to numerical values. In this paper we propose a general definition for continuous domains and another definition for partially ordered sets.
Vicenç Torra
ARES1
2009 Generation of Prototypes for Masking Sequences of Events
abstract
Sequences of categorical data are in common use to represent sequences of events. In order to transfer such data to third parties for their analysis, masking methods can be applied to satisfy privacy laws and avoid the disclosure of sensitive information. Masking methods distort the data so that privacy is kept at the expenses of some information loss. %Different methods exist, each one trying to find a good trade-off between the risk of disclosure and the information loss. Microaggregation is one of the existing masking methods. In microaggregation small clusters are automatically built and the values of the members of a cluster are substituted by the values of the prototype of that cluster. Due to the fact that microaggregation is an NP-hard problem, heuristic approaches have been developed. Existing methods are mainly devoted to numerical and categorical data. The extension of these methods to sequences of categorical data requires the definition of special algorithms for clustering and prototyping.Artificial Intelligence offers techniques and tools that are appropriate for symbolic data. As in our context the sequences are defined in terms of categorical (symbolic) values, such AI techniques are of special relevance. In this paper, we will use them to propose a new method for generating the prototype of a small group of sequences of categorical values. These results can later be used in e.g. microaggregation.
Aïda Valls, Cristina Gómez-Alonso, Vicenç Torra
ARES3
2009 Generation of synthetic data by means of fuzzy c-Regression
abstract
Problems related to data privacy are studied in the areas of privacy preserving data mining (PPDM) and statistical disclosure control (SDC). Their goal is to avoid the disclosure of sensitive or proprietary information to third parties. In this paper a new synthetic data generation method is proposed and the information loss and disclosure risk are measured. The method is based on fuzzy techniques. Informally, a fuzzy c-regression method is applied to the original data set and synthetic data is released with an appropriate information loss and disclosure risk depending on c. As other data protection methods do, our synthetic data generation procedure allows third parties to do some statistical computations with a limited risk of disclosure. The trade-off between data utility and data safety of our proposed method will be assessed.
Isaac Cano, Vicenç Torra
FUZZ-IEEE2
2009 On hesitant fuzzy sets and decision
abstract
Intuitionistic fuzzy sets (IFS) are a generalization of fuzzy sets where the membership is an interval. That is, membership, instead of being a single value, is an interval. A large number of operations have been defined for this type of fuzzy sets, and several applications have been developed in the last years. In this paper we describe hesitant fuzzy sets. They are another generalization of fuzzy sets. Although similar in intention to IFS, some basic differences on their interpretation and on their operators exist. In this paper we review their definition, the main results and we present an extension principle, which permits to generalize existing operations on fuzzy sets to this new type of fuzzy sets. We also discuss their use in decision making.
Vicenç Torra, Yasuo Narukawa
FUZZ-IEEE1
2009 Multidimensional generalized fuzzy integral
Yasuo Narukawa, Vicenç Torra
Fuzzy Sets Syst.2
2009 On the WOWA operator and its interpolation function
abstract
The weighted ordered weighted averaging (WOWA) operator is one of the existing aggregation methods that can be used to fuse numerical data. The application of this operator to a set of data requires an interpolation function. In this paper, we present a few results about the sensitivity of the operator according to the interpolation method used. © 2009 Wiley Periodicals, Inc.
Vicenç Torra, Zhenbang Lv
Int. J. Intell. Syst.1
2009 Towards the evaluation of time series protection methods
Jordi Nin, Vicenç Torra
Inf. Sci.2
2008 A Critique of k-Anonymity and Some of Its Enhancements
abstract
k-Anonymity is a privacy property requiring that all combinations of key attributes in a database be repeated at least for k records. It has been shown that k-anonymity alone does not always ensure privacy. A number of sophistications of k-anonymity have been proposed, like p-sensitive k-anonymity, l-diversity and t-closeness. This paper explores the shortcomings of those properties, none of which turns out to be completely convincing.
Josep Domingo-Ferrer, Vicenç Torra
ARES2
2008 Cluster-Specific Information Loss Measures in Data Privacy: A Review
abstract
Data protection mechanisms need to find a trade-off between information loss and disclosure risk. To this end, information loss and disclosure risk measures have been developed. Due to the fact that when data is published it is usual to ignore which kind of analyses a user will pursue with the data, generic information loss measures are used to analyse the impact of the perturbation method onto the data. Such generic information loss measures are defined in terms of a few general-enough statistics. Nevertheless, a more fine-grained analysis is needed for particular data uses. In this paper we provide the reader with a review of a few results on cluster-specific information loss measures. More specifically, we consider the case of using fuzzy clustering to the perturbated data.
Vicenç Torra, Susana Ladra
ARES1
2008 Domain extension for multidimensional generalized fuzzy integrals
abstract
We have recently studied multidimensional fuzzy integrals, partly motivated by the definition of citation indices. In this paper, we further study these integrals and consider the problem of domain extension. This problem arises when we consider the definition of a measure on the product space from two measures. We also discuss how these results are applied to the definition of citation indices.
Yasuo Narukawa, Vicenç Torra
FUZZ-IEEE2
2008 On intuitionistic fuzzy clustering for its application to privacy
abstract
Motivated by our research on specific information loss measures (in privacy preserving data mining) and our need to compare fuzzy clusters, we proposed in a recent paper a definition for intuitionistic fuzzy partitions. We showed how to define them in the framework of fuzzy clustering. That is, we introduced a method to define intuitionistic fuzzy partitions from the results of fuzzy clustering. In this paper we further study such intuitionistic fuzzy partitions and we extend our previous results with other types of fuzzy clustering algorithms.
Vicenç Torra, Sadaaki Miyamoto, Yasunori Endo, Josep Domingo-Ferrer
FUZZ-IEEE1
2008 Choquet Stieltjes Integral, Losonczi's Means and OWA Operators
Vicenç Torra, Yasuo Narukawa
MDAI1
2008 Towards a More Realistic Disclosure Risk Assessment
Jordi Nin, Javier Herranz, Vicenç Torra
Privacy in Statistical Databases3
2008 Rethinking rank swapping to decrease disclosure risk
Jordi Nin, Javier Herranz, Vicenç Torra
Data Knowl. Eng.3
2008 On the disclosure risk of multivariate microaggregation
Jordi Nin, Javier Herranz, Vicenç Torra
Data Knowl. Eng.3
2008 Measuring simultaneous belongingness for sets of objects
abstract
This paper presents the so-called measures of simultaneous belongingness. These measures are used in clustering for establishing in which extent two objects belong to the same clusters. In the case of fuzzy clustering, the measure also takes into account fuzzy membership. In this paper, we establish a more general framework and, in particular, we introduce a definition that permits to compute this measure for sets of objects (instead of only to pairs of them). © 2008 Wiley Periodicals, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
2008 Record linkage for database integration using fuzzy integrals
abstract
Given two-data databases, record linkage algorithms try to establish which records of these files contain information on the same individual. Standard record linkage algorithms assume that both files are described using the same attributes. In this article, we study the nonstandard case when the attributes are not the same. We apply aggregation operators for extracting relevant information for this purpose. We restrict to the case of numerical databases. © 2008 Wiley Periodicals, Inc.
Vicenç Torra, Jordi Nin
Int. J. Intell. Syst.1
2008 Editorial: Modeling decisions for artificial intelligence
Vicenç Torra, Yasuo Narukawa, Toho Gakuen
Int. J. Intell. Syst.1
2008 On the Comparison of Generic Information Loss Measures and Cluster-Specific Ones
abstract
Masking methods are to protect data bases prior to their public release. They mask an original data file so that the new file ensures the privacy of data respondents. Information loss measures have been developed to evaluate in which extent the masked file diverges from the corresponding original file, and in what extent the same analyses on both files lead to the same results. Generic information loss measures ignore the intended data use of the file. These are the standard measures when data has to be released (e.g. published in the web) and there is no control on what kind of analyses users would perform. In this paper we study generic information loss measures, and we compare such measures with respect to cluster-specific ones. That is, measures specifically defined for the case in which the user will do clustering with the original data. To do so, we define such measures and then we do an extensive comparison of the two measures. The paper shows that the generic measures can cope with the information loss related to clustering.
Susana Ladra, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2008 How to Group Attributes in Multivariate Microaggregation
abstract
Microaggregation is one of the most employed microdata protection methods. It builds clusters of at least k original records, and then replaces these records with the centroid of the cluster. When the number of attributes of the dataset is large, one usually splits the dataset into smaller blocks of attributes, and then applies microaggregation to each block, successively and independently. In this way, the effect of the noise introduced by microaggregation is reduced, at the cost of losing the k-anonymity property. In this work we show that, besides the specific microaggregation method, the value of the parameter k and the number of blocks in which the dataset is split, there exists another factor which influences the quality of the microaggregation: the way in which the attributes are grouped to form the blocks. When correlated attributes are grouped in the same block, the statistical utility of the protected dataset is higher. In contrast, when correlated attributes are dispersed into different blocks, the achieved anonymity is higher, and so, the disclosure risk is lower. We present quantitative evaluations of such statements based on different experiments on real datasets.
Jordi Nin, Javier Herranz, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2008 Guest Editors' Introduction
Vicenç Torra, Yasuo Narukawa
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2008 The h-Index and the Number of Citations: Two Fuzzy Integrals
abstract
In this paper, we review two of the most well-known citation indexes and establish their connections with the Choquet and Sugeno integrals. In particular, we show that the recently established h-index is a particular case of the Sugeno integral, and that the number of citations corresponds to the Choquet integral. In both cases, they use the same fuzzy measure. The results presented here permit one to envision new indexes defined in terms of fuzzy integrals using other types of fuzzy measures. A few considerations in this respect are also included in this paper. Indexes for taking into account recent research and the publisher credibility are outlined.
Vicenç Torra, Yasuo Narukawa
IEEE Trans. Fuzzy Syst.1
2007 Multidimensional Fuzzy Integrals
Yasuo Narukawa, Vicenç Torra
MDAI2
2007 A Multicriteria Fuzzy System Using Residuated Implication Operators and Fuzzy Arithmetic
Sandra A. Sandri, Christophe Sibertin-Blanc, Vicenç Torra
MDAI3
2007 Advances in smart cards
Josep Domingo-Ferrer, Joachim Posegga, Francesc Sebé, Vicenç Torra
Comput. Networks4
2007 Aggregation operators
Vicenç Torra, Yasuo Narukawa
Int. J. Approx. Reason.1
2007 Fuzzy measures and integrals in evaluation of strategies
Yasuo Narukawa, Vicenç Torra
Inf. Sci.2
2007 A View of Averaging Aggregation Operators
abstract
Aggregation operators have been used and studied for a long time. More than ten different types of means (including the arithmetic, geometric, and harmonic ones) were already studied 2000 years ago. Nevertheless, this topic has gained relevance in recent years. This is partly due to the increasing need of methods and operators for fusing information within computer programs. In this paper, we will present a personal view of the field, focusing on the application issues and on averaging aggregation operators. We will point out some research topics and open lines for future research.
Vicenç Torra, Yasuo Narukawa
IEEE Trans. Fuzzy Syst.1
2006 Establishing a benchmark for re-identification methods and its validation using fuzzy clustering
abstract
Privacy preserving data mining and statistical disclosure control are related fields with increasing importance nowadays. They aim is to allow the publication of sensible data without compromising the privacy of data respondents. To that end, masking methods have been designed so that data are distorted in a way that preserves confidentiality and data utility. Alternatively, methods have been constructed to generate synthetic data that have properties similar to the ones of the original data. At the same time, recent research in re-identification methods (record and variable matching) has been pushed forward due to the current interest on security issues and the huge amount of data stored in databases. However, there is no standard methodology for comparing alternative re-identification methods. In this paper we propose the use of masking methods and synthetic data generators for building benchmarks for matching methods. We validate our approach using fuzzy clustering.
Vicenç Torra, Josep Domingo-Ferrer
FUZZ-IEEE1
2006 Non-monotonic Fuzzy Measures and Intuitionistic Fuzzy Sets
Yasuo Narukawa, Vicenç Torra
MDAI2
2006 New Approach to the Re-identification Problem Using Neural Networks
Jordi Nin, Vicenç Torra
MDAI2
2006 On the Use of Variable-Size Fuzzy Clustering for Classification
Vicenç Torra, Sadaaki Miyamoto
MDAI1
2006 Distance Based Re-identification for Time Series, Analysis of Distances
Jordi Nin, Vicenç Torra
Privacy in Statistical Databases2
2006 Using Mahalanobis Distance-Based Record Linkage for Disclosure Risk Assessment
Vicenç Torra, John M. Abowd, Josep Domingo-Ferrer
Privacy in Statistical Databases1
2006 Generalized transformed t-conorm integral and multifold integral
Yasuo Narukawa, Vicenç Torra
Fuzzy Sets Syst.2
2006 Aggregation operators and decision modeling
Vicenç Torra, Yasuo Narukawa
Int. J. Approx. Reason.1
2006 The interpretation of fuzzy integrals and their application to fuzzy systems
Vicenç Torra, Yasuo Narukawa
Int. J. Approx. Reason.1
2006 Editorial
Vicenç Torra, Yasuo Narukawa, Sadaaki Miyamoto
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2006 Regression for ordinal variables without underlying continuous variables
Vicenç Torra, Josep Domingo-Ferrer, Josep Maria Mateo-Sanz, Michael Kwok-Po Ng
Inf. Sci.1
2006 Image clustering for the exploration of video sequences
abstract
Abstract In this article we present a system for the exploration of video sequences. The system, GAMBAL for the Exploration of Video Sequences (GAMBAL‐EVS), segments video sequences, extracting an image for each shot, and then clusters such images and presents them in a visualization system. The system allows the user to find similarities between images and to proceed through the video sequences to find the relevant ones.
Vicenç Torra, Sergi Lanau, Sadaaki Miyamoto
J. Assoc. Inf. Sci. Technol.1
2005 Fuzzy c-means for Fuzzy Hierarchical Clustering
abstract
This paper describes an algorithm for building fuzzy hierarchies. These are hierarchies where the elements can have fuzzy membership to the nodes. The paper presents an approach that mainly follows a bottom-up strategy, and describes the functions needed to operate with fuzzy variables. An example of the application of the approach is also presented
Vicenç Torra
FUZZ-IEEE1
2005 Modeling Decisions for Artificial Intelligence: Theory, Tools and Applications
Vicenç Torra, Yasuo Narukawa, Sadaaki Miyamoto
MDAI1
2005 Privacy in Data Mining
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.2
2005 Ordinal, Continuous and Heterogeneous k-Anonymity Through Microaggregation
Josep Domingo-Ferrer, Vicenç Torra
Data Min. Knowl. Discov.2
2005 Aggregation operators and models
Vicenç Torra
Fuzzy Sets Syst.1
2005 On the meta-knowledge choquet integral and related models
abstract
Choquet integral and multistep Choquet integrals have been used in recent years as models for decision making and information aggregation. Such models can be used to fuse information when information sources are not independent. A basic property of such models is that their output is monotonically increasing with respect to inputs. In this article we study two alternative models built on the basis of such Choquet integrals. The motivation is, on the one hand, to study the modeling capabilities of such operators and, on the other, to build models that are capable of approximating any arbitrary functions (not only monotonic ones). In this article we describe and study two models that are universal approximators. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 1017–1036, 2005.
Vicenç Torra, Yasuo Narukawa
Int. J. Intell. Syst.1
2005 Graphical Interpretation of the Twofold Integral and its Generalization
abstract
In this work we define a generalization of the twofold integral (generalization of the Choquet and Sugeno integrals) that operates in an arbitrary universal set. A graphical representation of this integral is also introduced.
Yasuo Narukawa, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2005 Exploration of textual document archives using a fuzzy hierarchical clustering algorithm in the GAMBAL system
Vicenç Torra, Sadaaki Miyamoto, Sergi Lanau
Inf. Process. Manag.1
2005 Fuzzy Measure and Probability Distributions: Distorted Probabilities
abstract
This work studies fuzzy measures and their application to data modeling. We focus on the particular case when fuzzy measures are distorted probabilities. We analyze their properties and introduce a new family of measures (m-dimensional distorted probabilities). The work finishes with the application of two dimensional distorted probabilities to a modeling problem. Results of the application of such fuzzy measure to data modeling using Choquet integral are discussed
Yasuo Narukawa, Vicenç Torra
IEEE Trans. Fuzzy Syst.2
2004 Object Positioning Based on Partial Preferences
Josep Maria Mateo-Sanz, Josep Domingo-Ferrer, Vicenç Torra
MDAI3
2004 On the Interpretation of Some Fuzzy Integrals
Vicenç Torra, Yasuo Narukawa
MDAI1
2004 OWA operators in data modeling and reidentification
abstract
This paper is devoted to the application of aggregation operators and to the application of ordered weighting averaging (OWA) operators to data mining. In particular, we consider two application of OWA operators in this field: model building and information extraction. The latter application is oriented to the reidentification procedures.
Vicenç Torra
IEEE Trans. Fuzzy Syst.1
2003 A multi-agent system approach for monitoring the prescription of restricted use antibiotics
Lluís Godo, Josep Puyol-Gruart, Jordi Sabater-Mir, Vicenç Torra, P. Barrufet, X. Fàbregas
Artif. Intell. Medicine4
2003 Median-based aggregation operators for prototype construction in ordinal scales
abstract
This article studies aggregation operators in ordinal scales for their application to clustering (more specifically, to microaggregation for statistical disclosure risk). In particular, we consider these operators in the process of prototype construction. This study analyzes main aggregation operators for ordinal scales [plurality rule, medians, Sugeno integrals (SI), and ordinal weighted means (OWM), among others] and shows the difficulties for their application in this particular setting. Then, we propose two approaches to solve the drawbacks and we study their properties. Special emphasis is given to the study of monotonicity because the operator is proven nonsatisfactory for this property. Exhaustive empirical work shows that in most practical situations, this cannot be considered a problem. © 2003 Wiley Periodicals, Inc.
Josep Domingo-Ferrer, Vicenç Torra
Int. J. Intell. Syst.2
2003 Semantic-based aggregation for statistical disclosure control
abstract
In this paper we show how clustering can be used to aggregate different versions of the same data set in order to discover confidential information. Having these tools helps to not publish data that could be reidentified, which is known as Statistical Disclosure Control. In particular, the paper is focused on the case of dealing with categorical values. © 2003 Wiley Periodicals, Inc.
Aïda Valls, Vicenç Torra, Josep Domingo-Ferrer
Int. J. Intell. Syst.2
2003 On the connections between statistical disclosure control for microdata and some artificial intelligence tools
Josep Domingo-Ferrer, Vicenç Torra
Inf. Sci.2
2002 Information-Theoretic Disclosure Risk Measures in Statistical Disclosure Control of Tabular Data
abstract
Statistical database protection is a part of information security which tries to prevent published statistical information (tables, individual records) from disclosing the contribution of specific respondents. This paper shows how to use information-theoretic concepts to measure disclosure risk for tabular data. The proposed disclosure risk measure is compatible with a broad class of disclosure protection methods and can be extended for computing disclosure risk for a set of linked tables.
Josep Domingo-Ferrer, Anna Oganian, Vicenç Torra
SSDBM3
2002 Special issue on hierarchical fuzzy systems
Vicenç Torra
Int. J. Intell. Syst.1
2002 A review of the construction of hierarchical fuzzy systems
abstract
Fuzzy rule-based systems are nowadays one of the most successful applications of fuzzy sets and fuzzy logic. Most applications use a flat set of fuzzy rules. However, in complex applications with a large set of variables, it is not appropriate to define the system with a flat set of rules because, among other problems, the number of rules increases exponentially with the number of variables. Hierarchical fuzzy systems are one of the alternatives presented in the literature to overcome this problem. In this article we review the latest results related with this type of fuzzy system. © 2002 Wiley Periodicals, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
2002 A Critique of the Sensitivity Rules Usually Employed for Statistical Table Protection
abstract
In statistical disclosure control of tabular data, sensitivity rules are commonly used to decide whether a table cell is sensitive and should therefore not be published. The most popular sensitivity rules are the dominance rule, the p%-rule and the pq-rule. The dominance rule has received critiques based on specific numerical examples and is being gradually abandoned by leading statistical agencies. In this paper, we construct general counterexamples which show that none of the above rules does adequately reflect disclosure risk if cell contributors or coalitions of them behave as intruders: in that case, releasing a cell declared non-sensitive can imply higher disclosure risk than releasing a cell declared sensitive. As possible solutions, we propose an alternative sensitivity rule based on the concentration of relative contributions. More generally, we suggest to complement a priori risk assessment based on sensitivity rules with a posteriori risk assessment which takes into account tables after they have been protected.
Josep Domingo-Ferrer, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2002 Editorial: Trends in Aggregation and Security Assessment for Inference Control in Statistical Databases
abstract
As e-commerce and Internet-based data handling become pervasive, companies and statistical agencies have the need to exploit the data they accumulate without violating citizens' privacy. Inference control is a discipline whose goal is to prevent published/exchanged data from being linked with the individual respondents they originated from. This special issue illustrates that inference control largely draws on soft computing and artificial intelligence techniques.
Vicenç Torra, Josep Domingo-Ferrer
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2002 Hierarchical Spherical Clustering
abstract
This work introduces an alternative representation for large dimensional data sets. Instead of using 2D or 3D representations, data is located on the surface of a sphere. Together with this representation, a hierarchical clustering algorithm is defined to analyse and extract the structure of the data. The algorithm builds a hierarchical structure (a dendrogram) in such a way that different cuts of the structure lead to different partitions of the surface of the sphere. This can be seen as a set of concentric spheres, each one being of different granularity. Also, to obtain an initial assignment of the data on the surface of the sphere, a method based on Sammon's mapping has been developed.
Vicenç Torra, Sadaaki Miyamoto
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
2002 Learning weights for the quasi-weighted means
abstract
We study the determination of weights for quasi-weighted means (also called quasi-linear means) when a set of examples is given. We consider first a simple case, the learning of weights for weighted means, and then we extend the approach to the more general case of a quasi-weighted mean. We consider the case of a known arbitrary generator f. The paper finishes considering the use of parametric functions that are suitable when the values to aggregate are measure values or ratio.
Vicenç Torra
IEEE Trans. Fuzzy Syst.1
2001 On Some Extensions for S-decomposable Fuzzy Measures
abstract
Practical applications of fuzzy integrals (Choquet and Sugeno integrals) require the definition of a fuzzy measure. However, when they are defined in a domain of n elements, this definition requires 2/sup n/ parameters. To avoid this large number of parameters, several families of measures have been introduced in the literature. In this work we present a new family that generalises some fuzzy measures but restricted so that the number of parameters is not large.
Vicenç Torra
FUZZ-IEEE1
2001 Author's reply
Vicenç Torra
Fuzzy Sets Syst.1
2001 A comparison of active set method and genetic algorithm approaches for learning weighting vectors in some aggregation operators
abstract
In this article we compare two contrasting methods, active set method (ASM) and genetic algorithms, for learning the weights in aggregation operators, such as weighted mean (WM), ordered weighted average (OWA), and weighted ordered weighted average (WOWA). We give the formal definitions for each of the aggregation operators, explain the two learning methods, give results of processing for each of the methods and operators with simple test datasets, and contrast the approaches and results. © 2001 John Wiley & Sons, Inc.
David F. Nettleton, Vicenç Torra
Int. J. Intell. Syst.2
2001 Aggregation of linguistic labels when semantics is based on antonyms
abstract
In this work, we introduce aggregation operators for linguistic labels (this is, ordinal scales) when different experts (or information sources) use different domains to express their knowledge. The aggregated value is computed (i) building first a unified framework, (ii) transforming all the initial values into this new framework, (iii) aggregating the transformed values, and (iv) finally applying a reversal transformation. Transformations and all the constructions are based on assuming an existing semantics for all the domains. In this work, we consider the semantics based on the existence of an antonym (or a set of them) for each element in the domain. This is equivalent to a semantics based on negation functions. © 2001 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
2001 Extending Choquet Integrals for Aggregation of Ordinal Values
Lluís Godo, Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2001 Fuzzy configuration of matching runtime implementation strategies
Angela C. Sodan, Vicenç Torra
Soft Comput.2
2000 Towards the Re-identification of Individuals in Data Files with Non-common Variables
Vicenç Torra
ECAI1
2000 Knowledge-based validation: Synthesis of diagnoses through synthesis of relations
Vicenç Torra
Fuzzy Sets Syst.1
2000 The WOWA operator and the interpolation function W*: Chen and Otto's interpolation method revisited
Vicenç Torra
Fuzzy Sets Syst.1
2000 Using classification as an aggregation tool in MCDM
Aïda Valls, Vicenç Torra
Fuzzy Sets Syst.2
2000 Introduction: The first Catalan conference on artificial intelligence
Vicenç Torra
Int. J. Intell. Syst.1
2000 On aggregation operators for ordinal qualitative information
abstract
In many fuzzy systems applications, values to be aggregated are of a qualitative nature. In that case, if one wants to compute some type of average, the most common procedure is to perform a numerical interpretation of the values, and then apply one of the well-known (the most suitable) numerical aggregation operators. However, if one wants to stick to a purely qualitative setting, choices are reduced to either weighted versions of max-min combinations or to a few existing proposals of qualitative versions of ordered weighted average (OWA) operators. In this paper, we explore the feasibility of defining a qualitative counterpart of the weighted mean operator without having to use necessarily any numerical interpretation of the values. We propose a method to average qualitative values, belonging to a (finite) ordinal scale, weighted with natural numbers, and based on the use of finite t-norms and t-conorms defined on the scale of values. Extensions of the method for other OWA-like and Choquet integral-type aggregations are also considered.
Lluís Godo, Vicenç Torra
IEEE Trans. Fuzzy Syst.2
1999 A multi-stage system in compilation environments
Vicenç Torra, Angela C. Sodan
Fuzzy Sets Syst.1
1999 Interpreting membership functions: A constructive approach
Vicenç Torra
Int. J. Approx. Reason.1
1999 On hierarchically S-decomposable fuzzy measures
abstract
In this work we introduce hierarchically decomposable fuzzy measures that allow the user to define measures more general than t-decomposable ones without having to deal with exponential complexity. ©1999 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
1999 On some relationships between hierarchies of quasiarithmetic means and neural networks
abstract
In this work, we establish the relations between neural networks and hierarchies of quasiarithmetic means. We show that a neural network with the same activation function in all the neurons gives an output that is isomorphic to the result that can be obtained with a hierarchy of quasiarithmetic means. From this result, we show that hierarchies of quasiarithmetic means are universal approximations. ©1999 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
1999 On the semantics of qualitative attributes in knowledge elicitation
abstract
In this work we propose a new discretization method for quantitative domains. We give an overview of the methods in the literature and we introduce a new one that returns better results. As the need for a discretization method appeared in the definition of a procedure for evaluating the semantics of linguistic labels defined in a previous paper, we also give in this paper an overview of this semantics and of its analysis. ©1999 John Wiley & Sons, Inc.
Aïda Valls, Vicenç Torra
Int. J. Intell. Syst.2
1998 On Considering Constraints of Different Importance in Fuzzy Constraint Satisfaction Problems
abstract
Several real-world applications (e.g., scheduling, configuration, …) can be formulated as Constraint Satisfaction Problems (CSP). In these cases, a set of variables have to be settled to a value with the requirement that they satisfy a set of constraints. Classical CSPs are defined only by means of crisp (Boolean) constraints. However, as sometimes Boolean constraints are too strict in relation to human reasoning, fuzzy constraints were introduced. When fuzzy constraints are considered, human reasoning usually performs some compensation between alternatives. Thus other operators than t-norms are advisable. Besides of that, not all constraints can be considered with equal importance. In this paper we show that the WOWA operator can consider both aspects: compensation between constraints and constraints of different importance.
Vicenç Torra
Int. J. Uncertain. Fuzziness Knowl. Based Syst.1
1997 Synthesis of fuzzy relations for knowledge based systems validation
Vicenç Torra
Fuzzy Sets Syst.1
1997 The weighted OWA operator
abstract
One of the properties that the OWA operator satisfies is commutativity. This condition, that is not satisfied by the weighted mean, stands for equal reliability of all the information sources that supply the data. In this article we define a new combination function, the WOWA (Weighted OWA), that combines the advantages of the OWA operator and the ones of the weighted mean. We study some of its properties and show how it can be extended to deal with linguistic labels. © 1997 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
1996 Negation functions based semantics for ordered linguistic labels
abstract
After arguing that in the knowledge acquisition framework experts cannot always supply a precise semantics for the linguistic labels they use, we show that negation functions over an ordered set of linguistics labels induce a semantics. We study the semantics induced by classical negation, functions from L to L, and also the one induced by negation functions from L to parts of L. © 1996 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
1995 Combining fuzzy sets: The geometric consensus function family
Vicenç Torra
Fuzzy Sets Syst.1
1995 A new combination function in evidence theory
abstract
This article studies the combination of basic probability assignment (bpa) in evidence theory. After introducing an interpretation of the mass function we show that given two bpa Dempster's rule of combination does not build a coherent bpa with respect to the interpretation. Next we give a new combination function that overcomes this problem and study some of its properties. © 1995 John Wiley & Sons, Inc.
Vicenç Torra
Int. J. Intell. Syst.1
1995 Towards an automatic consensus generator tool: EGAC
abstract
Automatic knowledge acquisition for expert systems has attracted much attention. Several algorithms that infer concept descriptions from a given set of training examples have been developed to aid in this task, some of them elicit concepts from examples organized in data matrices. These algorithms infer from different training examples (or different matrices defined by a set of experts) slightly different concept descriptions. Herein we propose a method based on synthesis of judgements, fuzzy sets and classification methods that applied to a set of data matrices builds an agreed one that synthesises the information contained in the set of matrices. The method proposed can be applied to data matrices with attributes of several types: measure and ratio quantitative attributes, and ordered and nonordered qualitative ones.>
Vicenç Torra, Ulises Cortés
IEEE Trans. Syst. Man Cybern.1