Ruggero G. Pensa

dblp:23/1886 · also Ruggero Gaetano Pensa · DBLP profile ↗
← Back
17ranked-venue papers in the field
5as first author
4since 2021 · last 2026
0000-0001-5145-3438ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 16 (5 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 Differentiable parameter-less co-clustering using graph neural networks
abstract
Abstract Co-clustering refers to the simultaneous clustering of rows and columns in a data matrix, uncovering joint patterns between two distinct sets, such as documents and terms or users and products. Traditional co-clustering algorithms typically rely on discrete optimization techniques based on enumeration, which can limit both scalability and flexibility. In this paper, we introduce a differentiable programming approach to co-clustering that enables the continuous optimization of co-partitions using graph neural networks. Our method is grounded in an associative co-clustering quality measure that is independent of the number of clusters and dynamically adjusts this parameter by jointly considering both partitions. By leveraging automatic differentiation and graph neural networks, our approach scales to very large datasets while maintaining high-quality co-cluster structures. We evaluate our method using different types of graph neural networks and initialization strategies. Furthermore, when compared with recent state-of-the-art methods for co-clustering and graph clustering, our approach achieves competitive or superior results in terms of accuracy. Most importantly, it is the only algorithm that successfully completes on the largest benchmark dataset.
Alessio Ragno, Pierre-Angelo Peyrie, Marc Plantevit, Ruggero G. Pensa, Céline Robardet
Data Min. Knowl. Discov.4
2025 Fair Associative Co-clustering
Federico Peiretti, Ruggero G. Pensa
ECML/PKDD (1)2
2025 Differentially Private Associative Co-clustering
abstract
Co-clustering is a useful tool that extracts summary information from a data matrix in terms of row and column clusters, and gives a succinct representation of the data. However, if the matrix contains data about individuals, such representations could leak their privacy-sensitive information. In terms of privacy disclosure, co-clustering is even more harmful than clustering, because of the additional information carried by the column partition. However, to the best of our knowledge, the problem of privacy-preserving co-clustering has never been studied. To fill this gap, we consider a recent co-clustering algorithm, based on a de-normalized version of the Goodman-Kruskal’s τ association measure, which has a good property from a differential privacy perspective, and is supposed not to consume an excessive amount of privacy budget. This leads to a privacy-preserving co-clustering algorithm that satisfies the definition of differential privacy while providing good partitioning solutions. Our algorithm is based on a prototype-based optimization strategy that makes it fast and actionable in realistic privacy-preserving data management and analysis scenarios, as shown by our extensive experimental validation.
Elena Battaglia, Ruggero G. Pensa
SDM2
2021 Differentially Private Distance Learning in Categorical Data
abstract
Abstract Most privacy-preserving machine learning methods are designed around continuous or numeric data, but categorical attributes are common in many application scenarios, including clinical and health records, census and survey data. Distance-based methods, in particular, have limited applicability to categorical data, since they do not capture the complexity of the relationships among different values of a categorical attribute. Although distance learning algorithms exist for categorical data, they may disclose private information about individual records if applied to a secret dataset. To address this problem, we introduce a differentially private family of algorithms for learning distances between any pair of values of a categorical attribute according to the way they are co-distributed with the values of other categorical attributes forming the so-called context. We define different variants of our algorithm and we show empirically that our approach consumes little privacy budget while providing accurate distances, making it suitable in distance-based applications, such as clustering and classification.
Elena Battaglia, Simone Celano, Ruggero G. Pensa
Data Min. Knowl. Discov.3
2020 Towards Content Sensitivity Analysis
abstract
With the availability of user-generated content in the Web, malicious users dispose of huge repositories of private (and often sensitive) information regarding a large part of the world’s population. The self-disclosure of personal information, in the form of text, pictures and videos, exposes the authors of such contents (and not only them) to many criminal acts such as identity thefts, stalking, burglary, frauds, and so on. In this paper, we propose a way to evaluate the harmfulness of any form of content by defining a new data mining task called content sensitivity analysis . According to our definition, a score can be assigned to any object (text, picture, video...) according to its degree of sensitivity. Even though the task is similar to sentiment analysis, we show that it has its own peculiarities and may lead to a new branch of research. Thanks to some preliminary experiments, we show that content sensitivity analysis can not be addressed as a simple binary classification task.
Elena Battaglia, Livio Bioglio, Ruggero G. Pensa
IDA3
2017 TrAnET: Tracking and Analyzing the Evolution of Topics in Information Networks
Livio Bioglio, Ruggero G. Pensa, Valentina Rho
ECML/PKDD (3)2
2017 Shaping City Neighborhoods Leveraging Crowd Sensors
Giuseppe Rizzo 0002, Rosa Meo, Ruggero G. Pensa, Giacomo Falcone, Raphaël Troncy
Inf. Syst.3
2016 A centrality-based measure of user privacy in online social networks
abstract
The risks due to a global and unaware diffusion of our personal data cannot be overlooked when more than two billion people are estimated to be registered in at least one of the most popular online social networks. As a consequence, privacy has become a primary concern among social network analysts and Web/data scientists. Some studies propose to “measure” users' profile privacy according to their privacy settings but do not consider the topological properties of the social network adequately. In this paper, we address this limitation and define a centrality-based privacy score to measure the objective user privacy risk according to the network properties. We analyze the effectiveness of our measures on a large network of real Facebook users.
Ruggero G. Pensa, Gianpiero di Blasi
ASONAM1
2014 Hierarchical co-clustering: off-line and incremental approaches
Ruggero G. Pensa, Dino Ienco, Rosa Meo
Data Min. Knowl. Discov.1
2013 Parameter-less co-clustering for star-structured heterogeneous data
Dino Ienco, Céline Robardet, Ruggero G. Pensa, Rosa Meo
Data Min. Knowl. Discov.3
2012 From Context to Distance: Learning Dissimilarity for Categorical Data Clustering
abstract
Clustering data described by categorical attributes is a challenging task in data mining applications. Unlike numerical attributes, it is difficult to define a distance between pairs of values of a categorical attribute, since the values are not ordered. In this article, we propose a framework to learn a context-based distance for categorical attributes. The key intuition of this work is that the distance between two values of a categorical attribute A i can be determined by the way in which the values of the other attributes A j are distributed in the dataset objects: if they are similarly distributed in the groups of objects in correspondence of the distinct values of A i a low value of distance is obtained. We propose also a solution to the critical point of the choice of the attributes A j . We validate our approach by embedding our distance learning framework in a hierarchical clustering algorithm. We applied it on various real world and synthetic datasets, both low and high-dimensional. Experimental results show that our method is competitive with respect to the state of the art of categorical data clustering approaches. We also show that our approach is scalable and has a low impact on the overall computational time of a clustering task.
Dino Ienco, Ruggero G. Pensa, Rosa Meo
ACM Trans. Knowl. Discov. Data2
2009 Social Network Analysis as Knowledge Discovery Process: A Case Study on Digital Bibliography
abstract
Today digital bibliographies are a powerful instrument that collects a great amount of data about scientific publications. Digital bibliographies have been used as basis of many studies focused on the knowledge extraction in databases. Here we present anew methodology for mining knowledge in this field. Our approach aims to apply the potential of social network analysis techniques to accomplish this task, using a network representation of bibliography data. Besides we use some data mining techniques applied on social network representations in order to enrich this new point of view and to evolve our methodology towards a comprehensive local and global bibliography analysis workflow seen as a knowledge discovery process.
Michele Coscia, Fosca Giannotti, Ruggero G. Pensa
ASONAM3
2009 Context-Based Distance Learning for Categorical Data Clustering
Dino Ienco, Ruggero G. Pensa, Rosa Meo
IDA2
2009 Parameter-Free Hierarchical Co-clustering by n-Ary Splits
Dino Ienco, Ruggero G. Pensa, Rosa Meo
ECML/PKDD (1)2
2008 Constrained Co-clustering of Gene Expression Data
abstract
In many applications, the expert interpretation of co-clustering is easier than for mono-dimensional clustering. Co-clustering aims at computing a bi-partition that is a collection of co-clusters: each co-cluster is a group of objects associated to a group of attributes and these associations can support interpretations. Many constrained clustering algorithms have been proposed to exploit the domain knowledge and to improve partition relevancy in the mono-dimensional case (e.g., using the so-called must-link and cannot-link constraints). Here, we consider constrained co-clustering not only for extended must-link and cannot-link constraints (i.e., both objects and attributes can be involved), but also for interval constraints that enforce properties of co-clusters when considering ordered domains. We propose an iterative co-clustering algorithm which exploits user-defined constraints while minimizing the sum-squared residues, i.e., an objective function introduced for gene expression data clustering by Cho et al. (2004). We illustrate the added value of our approach in two applications on gene expression data.
Ruggero G. Pensa, Jean-François Boulicaut
SDM1
2005 From Local Pattern Mining to Relevant Bi-cluster Characterization
Ruggero G. Pensa, Jean-François Boulicaut
IDA1
2005 A Bi-clustering Framework for Categorical Data
Ruggero G. Pensa, Céline Robardet, Jean-François Boulicaut
PKDD1