Ruggero G. Pensa

dblp:23/1886 · also Ruggero Gaetano Pensa · DBLP profile ↗
← Back
42ranked-venue papers
13as first author
10since 2021 · last 2026
0000-0001-5145-3438ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 11 first-author · 7 since 2021Databases, data management, data science and information retrieval · 17 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Differentiable parameter-less co-clustering using graph neural networks
abstract
Abstract Co-clustering refers to the simultaneous clustering of rows and columns in a data matrix, uncovering joint patterns between two distinct sets, such as documents and terms or users and products. Traditional co-clustering algorithms typically rely on discrete optimization techniques based on enumeration, which can limit both scalability and flexibility. In this paper, we introduce a differentiable programming approach to co-clustering that enables the continuous optimization of co-partitions using graph neural networks. Our method is grounded in an associative co-clustering quality measure that is independent of the number of clusters and dynamically adjusts this parameter by jointly considering both partitions. By leveraging automatic differentiation and graph neural networks, our approach scales to very large datasets while maintaining high-quality co-cluster structures. We evaluate our method using different types of graph neural networks and initialization strategies. Furthermore, when compared with recent state-of-the-art methods for co-clustering and graph clustering, our approach achieves competitive or superior results in terms of accuracy. Most importantly, it is the only algorithm that successfully completes on the largest benchmark dataset.
Alessio Ragno, Pierre-Angelo Peyrie, Marc Plantevit, Ruggero G. Pensa, Céline Robardet
Data Min. Knowl. Discov.4
2025 Fair Associative Co-clustering
Federico Peiretti, Ruggero G. Pensa
ECML/PKDD (1)2
2025 Differentially Private Associative Co-clustering
abstract
Co-clustering is a useful tool that extracts summary information from a data matrix in terms of row and column clusters, and gives a succinct representation of the data. However, if the matrix contains data about individuals, such representations could leak their privacy-sensitive information. In terms of privacy disclosure, co-clustering is even more harmful than clustering, because of the additional information carried by the column partition. However, to the best of our knowledge, the problem of privacy-preserving co-clustering has never been studied. To fill this gap, we consider a recent co-clustering algorithm, based on a de-normalized version of the Goodman-Kruskal’s τ association measure, which has a good property from a differential privacy perspective, and is supposed not to consume an excessive amount of privacy budget. This leads to a privacy-preserving co-clustering algorithm that satisfies the definition of differential privacy while providing good partitioning solutions. Our algorithm is based on a prototype-based optimization strategy that makes it fast and actionable in realistic privacy-preserving data management and analysis scenarios, as shown by our extensive experimental validation.
Elena Battaglia, Ruggero G. Pensa
SDM2
2025 Explaining Random Forest and XGBoost with Shallow Decision Trees by Co-clustering Feature Importance
abstract
Abstract Transparency is a non-functional requirement of machine learning that promotes interpretable models and easily explainable outcomes. Unfortunately, interpretable classification models, such as linear, rule-based, and decision tree models, are superseded by more accurate but complex learning paradigms, such as deep neural networks and ensemble methods. More specifically, for tabular data classification, models based on tree ensembles, such as random forest or XGBoost, are still competitive compared to deep learning ones and are often preferred to the latter. However, they share the same interpretability issues, due to the complexity of the learned model and, consequently, offer low explainability of predictions. Existing solutions consist of computing some feature importance score or extracting an approximate surrogate model from the learned tree ensemble. However, these methods lead to surrogate models with either poor fidelity or questionable comprehensibility. In this paper, we propose to improve this trade-off using Goodman-Kruskal’s association measure to find groups of instances with predictions that are explained by shared groups of features. To build this structure, instances are first described by SHAP values, which capture local feature importance, and then co-clustered with features on the basis of these SHAP values. Next, a surrogate model is built as a set of shallow decision trees learned for the different groups of instances and subsets of relevant features. Our experiments show that our method produces surrogate models that explain random forest and XGBoost classifiers with competitive fidelity and higher comprehensibility compared to recent state-of-the-art competitors.
Ruggero G. Pensa, Anton Crombach, Sergio Peignier, Christophe Rigotti
Mach. Learn.1
2024 Combining SHAP-Driven Co-clustering and Shallow Decision Trees to Explain XGBoost
Ruggero G. Pensa, Anton Crombach, Sergio Peignier, Christophe Rigotti
DS (1)1
2024 Fast parameterless prototype-based co-clustering
abstract
Abstract Tensor co-clustering algorithms have been proven useful in many application scenarios, such as recommender systems, biological data analysis and the analysis of complex and evolving networks. However, they are significantly affected by wrong parameter configurations, since, at the very least, they require the cluster number to be set for each mode of the matrix/tensor, although they typically have other algorithm-specific hyper-parameters that need to be fine-tuned. Among the few known objective functions that can be optimized without setting these parameters, the Goodman–Kruskal $$\tau $$ τ —a statistical association measure that estimates the strength of the link between two or more discrete random variables—has proven its effectiveness in complex matrix and tensor co-clustering applications. However, its optimization in a co-clustering setting is tricky and, so far, has leaded to very slow and, at least in some specific but not unfrequent cases, inaccurate algorithms, due to its normalization term. In this paper, we investigate some interesting mathematical properties of $$\tau $$ τ , and propose a new simplified objective function with the ability of discovering an arbitrary and a priori unspecified number of good-quality co-clusters. Additionally, the new objective function definition allows for a novel prototype-based optimization strategy that enables the fast execution of matrix and higher-order tensor co-clustering. We show experimentally that the new algorithm preserves or even improves the quality of the discovered co-clusters by outperforming state-of-the-art competing approaches, while reducing the execution time by at least two orders of magnitude.
Elena Battaglia, Federico Peiretti, Ruggero G. Pensa
Mach. Learn.3
2023 A parameter-less algorithm for tensor co-clustering
abstract
Abstract The majority of the data produced by human activities and modern cyber-physical systems involve complex relations among their features. Such relations can be often represented by means of tensors, which can be viewed as generalization of matrices and, as such, can be analyzed by using higher-order extensions of existing machine learning methods, such as clustering and co-clustering. Tensor co-clustering, in particular, has been proven useful in many applications, due to its ability of coping with n-modal data and sparsity. However, setting up a co-clustering algorithm properly requires the specification of the desired number of clusters for each mode as input parameters. This choice is already difficult in relatively easy settings, like flat clustering on data matrices, but on tensors it could be even more frustrating. To face this issue, we propose a new tensor co-clustering algorithm that does not require the number of desired co-clusters as input, as it optimizes an objective function based on a measure of association across discrete random variables (called Goodman and Kruskal’s $$\tau$$ τ ) that is not affected by their cardinality. We introduce different optimization schemes and show their theoretical and empirical convergence properties. Additionally, we show the effectiveness of our algorithm on both synthetic and real-world datasets, also in comparison with state-of-the-art co-clustering methods based on tensor factorization and latent block models.
Elena Battaglia, Ruggero G. Pensa
Mach. Learn.2
2023 FairSwiRL: fair semi-supervised classification with representation learning
abstract
Abstract Semi-supervised learning has shown its potential in many real-world applications where only few labeled examples are available. However, when some fairness constraints need to be satisfied, semi-supervised classification models often struggle as they are required to cope with the lack of sufficient information for predicting the target variable while forgetting its relationships with any sensitive and potentially discriminatory attribute. To address this issue, we propose a fair semi-supervised representation learning architecture that leads to fair and accurate classification results even in very challenging scenarios with few labeled (but biased) instances. We show experimentally that our model can be easily adopted in very general settings, as the learned representations may be employed to train any supervised classifier. Moreover, when applied to several synthetic and real-world datasets, our method is competitive with state-of-the-art fair semi-supervised approaches.
Mattia Cerrato, Dino Ienco, Ruggero G. Pensa, Roberto Esposito
Mach. Learn.4
2021 Differentially Private Distance Learning in Categorical Data
abstract
Abstract Most privacy-preserving machine learning methods are designed around continuous or numeric data, but categorical attributes are common in many application scenarios, including clinical and health records, census and survey data. Distance-based methods, in particular, have limited applicability to categorical data, since they do not capture the complexity of the relationships among different values of a categorical attribute. Although distance learning algorithms exist for categorical data, they may disclose private information about individual records if applied to a secret dataset. To address this problem, we introduce a differentially private family of algorithms for learning distances between any pair of values of a categorical attribute according to the way they are co-distributed with the values of other categorical attributes forming the so-called context. We define different variants of our algorithm and we show empirically that our approach consumes little privacy budget while providing accurate distances, making it suitable in distance-based applications, such as clustering and classification.
Elena Battaglia, Simone Celano, Ruggero G. Pensa
Data Min. Knowl. Discov.3
2021 ESA☆: A generic framework for semi-supervised inductive learning
Dino Ienco, Roberto Esposito, Ruggero G. Pensa
Neurocomputing4
2020 Towards Content Sensitivity Analysis
abstract
With the availability of user-generated content in the Web, malicious users dispose of huge repositories of private (and often sensitive) information regarding a large part of the world’s population. The self-disclosure of personal information, in the form of text, pictures and videos, exposes the authors of such contents (and not only them) to many criminal acts such as identity thefts, stalking, burglary, frauds, and so on. In this paper, we propose a way to evaluate the harmfulness of any form of content by defining a new data mining task called content sensitivity analysis . According to our definition, a score can be assigned to any object (text, picture, video...) according to its degree of sensitivity. Even though the task is similar to sentiment analysis, we show that it has its own peculiarities and may lead to a new branch of research. Thanks to some preliminary experiments, we show that content sensitivity analysis can not be addressed as a simple binary classification task.
Elena Battaglia, Livio Bioglio, Ruggero G. Pensa
IDA3
2020 Ranking by inspiration: a network science approach
Livio Bioglio, Valentina Rho, Ruggero G. Pensa
Mach. Learn.3
2020 Enhancing Graph-Based Semisupervised Learning via Knowledge-Aware Data Embedding
abstract
Semisupervised learning (SSL) is a family of classification methods conceived to reduce the amount of required labeled information in the training phase. Graph-based methods are among the most popular semisupervised strategies: the nearest neighbor graph is built in such a way that the manifold of the data is captured and the labeled information is propagated to target samples along the structure of the manifold. Research in graph-based SSL has mainly focused on two aspects: 1) the construction of the k -nearest neighbors graph and/or 2) the propagation algorithm providing the classification. Differently from the previous literature, in this article, we focus on the data representation with the aim of incorporating semisupervision earlier in the process. To this end, we propose an algorithm that learns a new knowledge-aware data embedding via an ensemble of semisupervised autoencoders to enhance a graph-based semisupervised classification. The experiments carried out on different classification tasks demonstrate the benefit of our approach.
Dino Ienco, Ruggero G. Pensa
IEEE Trans. Neural Networks Learn. Syst.2
2019 Parameter-Less Tensor Co-clustering
Elena Battaglia, Ruggero G. Pensa
DS2
2019 Deep Triplet-Driven Semi-supervised Embedding Clustering
Dino Ienco, Ruggero G. Pensa
DS2
2018 Semi-Supervised Clustering With Multiresolution Autoencoders
abstract
In most real world clustering scenarios, experts generally dispose of limited background information, but such knowledge is valuable and may guide the analysis process. Semi-supervised clustering can be used to drive the algorithmic process with prior knowledge and to enable the discovery of clusters that meet the analyst's expectations. Usually, in the semi-supervised clustering setting, the background knowledge is converted to some kind of constraint and, successively, metric learning or constrained clustering are adopted to obtain the final data partition. Conversely, we propose a new semi-supervised clustering algorithm that directly exploits prior knowledge, under the form of labeled examples, avoiding the necessity to derive constraints. Our algorithm employs a multiresolution strategy to generate an ensemble of semi-supervised autoencoders that fit the data together with the background knowledge. Successively, the network models are employed to supply a new embedding representation on which clustering is performed. The proposed strategy is evaluated on a set of real-world benchmarks also in comparison with well-known state-of-the-art semi-supervised clustering methods. The experimental results highlight the benefit of directly leveraging the prior knowledge and show the quality of the representation learnt by the multiresolution schema.
Dino Ienco, Ruggero G. Pensa
IJCNN2
2017 Measuring the Inspiration Rate of Topics in Bibliographic Networks
Livio Bioglio, Valentina Rho, Ruggero G. Pensa
DS3
2017 Concept-Enhanced Multi-view Co-clustering of Document Data
Valentina Rho, Ruggero G. Pensa
ISMIS2
2017 TrAnET: Tracking and Analyzing the Evolution of Topics in Information Networks
Livio Bioglio, Ruggero G. Pensa, Valentina Rho
ECML/PKDD (3)2
2017 A privacy self-assessment framework for online social networks
Ruggero G. Pensa, Gianpiero di Blasi
Expert Syst. Appl.1
2017 Shaping City Neighborhoods Leveraging Crowd Sensors
Giuseppe Rizzo 0002, Rosa Meo, Ruggero G. Pensa, Giacomo Falcone, Raphaël Troncy
Inf. Syst.3
2017 Introduction to the special issue on dynamic networks and knowledge discovery
Céline Rouveirol, Ruggero G. Pensa, Rushed Kanawati
Mach. Learn.2
2017 A Semisupervised Approach to the Detection and Characterization of Outliers in Categorical Data
abstract
In this paper, we introduce a new approach of semisupervised anomaly detection that deals with categorical data. Given a training set of instances (all belonging to the normal class), we analyze the relationship among features for the extraction of a discriminative characterization of the anomalous instances. Our key idea is to build a model that characterizes the features of the normal instances and then use a set of distance-based techniques for the discrimination between the normal and the anomalous instances. We compare our approach with the state-of-the-art methods for semisupervised anomaly detection. We empirically show that a specifically designed technique for the management of the categorical data outperforms the general-purpose approaches. We also show that, in contrast with other approaches that are opaque because their decision cannot be easily understood, our proposed approach produces a discriminative model that can be easily interpreted and used for the exploration of the data.
Dino Ienco, Ruggero G. Pensa, Rosa Meo
IEEE Trans. Neural Networks Learn. Syst.2
2016 A centrality-based measure of user privacy in online social networks
abstract
The risks due to a global and unaware diffusion of our personal data cannot be overlooked when more than two billion people are estimated to be registered in at least one of the most popular online social networks. As a consequence, privacy has become a primary concern among social network analysts and Web/data scientists. Some studies propose to “measure” users' profile privacy according to their privacy settings but do not consider the topological properties of the social network adequately. In this paper, we address this limitation and define a centrality-based privacy score to measure the objective user privacy risk according to the network properties. We analyze the effectiveness of our measures on a large network of real Facebook users.
Ruggero G. Pensa, Gianpiero di Blasi
ASONAM1
2016 A Semi-supervised Approach to Measuring User Privacy in Online Social Networks
Ruggero G. Pensa, Gianpiero di Blasi
DS1
2016 Positive and unlabeled learning in categorical data
Dino Ienco, Ruggero G. Pensa
Neurocomputing2
2016 Recommending multimedia visiting paths in cultural heritage applications
Ilaria Bartolini, Vincenzo Moscato, Ruggero G. Pensa, Antonio Penta, Antonio Picariello, Carlo Sansone, Maria Luisa Sapino
Multim. Tools Appl.3
2014 Hierarchical co-clustering: off-line and incremental approaches
Ruggero G. Pensa, Dino Ienco, Rosa Meo
Data Min. Knowl. Discov.1
2014 Leveraging additional knowledge to support coherent bicluster discovery in gene expression data
abstract
The increasing availability of gene expression data has encouraged the development of purposely-built intelligent data analysis techniques. Grouping genes characterized by similar expression patterns is a widely accepted – and often mandatory – analysis step. Despite the fact that a number of biclu stering methods have been developed to discover clusters of genes exhibiting a similar expression profile under a subgroup of experimental conditions, approaches driven by similarity measures based on expression profiles alone may lead to groups that are biologically meaningless. The integration of additional information, such as functional annotations, into biclustering algorithms can instead provide an effective support for identifying meaningful gene associations. In this paper we propose a new biclustering approach called Additional Information Driven Iterative Signature Algorithm, AID-ISA. It supports the extraction of biologically relevant biclusters by leveraging additional knowledge. We show that AID-ISA allows the discovery of coherent biclusters in baker's yeast and human gene expression data sets.
Alessia Visconti, Francesca Cordero, Ruggero G. Pensa
Intell. Data Anal.3
2013 Parameter-less co-clustering for star-structured heterogeneous data
Dino Ienco, Céline Robardet, Ruggero G. Pensa, Rosa Meo
Data Min. Knowl. Discov.3
2013 Guest Editorial
abstract
Modeling and analyzing networks is a major emerging topic in different research areas, such as computational biology, social science, document retrieval and social web applications.By connecting objects, it is possible to obtain an intuitive and global view of the relationships among components of a complex system.Nowadays, scientific communities have access to huge volume of network-structured data, such as social networks, gene/proteins/metabolic networks, sensor networks, and peer-to-peer networks.Often, data is collected at different time points allowing capturing a dynamic trend of the observed network.Consequently, the time component plays a key role in the comprehension of the evolutionary behavior of the studied network (evolution of the network structure and/or of flows within the system).Time can help to determine the real causal relationships within, for instance, gene activations, link creation, and information flow.Handling such data is a major challenge for current research in machine learning and data mining, and it has led to the development of recent innovative techniques that consider complex/multi-level networks, time-evolving graphs, heterogeneous information (nodes and links), and requires scalable algorithms that are able to manage large-scale complex networks.This special issue is the follow-up of the Dynamic Networks and Knowledge Discovery workshop (DyNaK) 1 that has been held in conjunction to ECML-PKDD 2011 at Barcelona on September 24th 2011.The workshop was motivated by the interest of providing a meeting point for scientists with different backgrounds who are interested in the study of large-scale dynamic complex networks.The workshop has attracted 18 submissions out of which 9 papers has been accepted.The workshop has gathered more than 30 participants and was also the host of three highly appreciated invited keynotes and one industrial talk.Building on the success of the DyNaK workshop, an open call for papers has been issued for this special issue, focusing on the major topic discussed in the workshop: analyzing, modeling and mining large-scale real network.15 high quality papers have been received; each of which has been reviewed by three reviewers.Only 7 contributions were finally selected.These contributions show the vitality of the field: a broad panel of techniques are applied to modeling the dynamics of complex systems, using a wide set of formalisms ranging from descriptive rules to Probabilistic Real-Time Automata.Application fields are also wide: vision, opinion diffusion in social network, business process modeling and text mining.In Internal link prediction: a new approach for predicting links in bipartite graphs, Allali et al. present an algorithm for predicting internal link in bipartite graph.They address the problem of predicting
Ruggero G. Pensa, Francesca Cordero, Céline Rouveirol, Rushed Kanawati
Intell. Data Anal.1
2012 From Context to Distance: Learning Dissimilarity for Categorical Data Clustering
abstract
Clustering data described by categorical attributes is a challenging task in data mining applications. Unlike numerical attributes, it is difficult to define a distance between pairs of values of a categorical attribute, since the values are not ordered. In this article, we propose a framework to learn a context-based distance for categorical attributes. The key intuition of this work is that the distance between two values of a categorical attribute A i can be determined by the way in which the values of the other attributes A j are distributed in the dataset objects: if they are similarly distributed in the groups of objects in correspondence of the distinct values of A i a low value of distance is obtained. We propose also a solution to the critical point of the choice of the attributes A j . We validate our approach by embedding our distance learning framework in a hierarchical clustering algorithm. We applied it on various real world and synthetic datasets, both low and high-dimensional. Experimental results show that our method is competitive with respect to the state of the art of categorical data clustering approaches. We also show that our approach is scalable and has a low impact on the overall computational time of a clustering task.
Dino Ienco, Ruggero G. Pensa, Rosa Meo
ACM Trans. Knowl. Discov. Data2
2009 Social Network Analysis as Knowledge Discovery Process: A Case Study on Digital Bibliography
abstract
Today digital bibliographies are a powerful instrument that collects a great amount of data about scientific publications. Digital bibliographies have been used as basis of many studies focused on the knowledge extraction in databases. Here we present anew methodology for mining knowledge in this field. Our approach aims to apply the potential of social network analysis techniques to accomplish this task, using a network representation of bibliography data. Besides we use some data mining techniques applied on social network representations in order to enrich this new point of view and to evolve our methodology towards a comprehensive local and global bibliography analysis workflow seen as a knowledge discovery process.
Michele Coscia, Fosca Giannotti, Ruggero G. Pensa
ASONAM3
2009 Context-Based Distance Learning for Categorical Data Clustering
Dino Ienco, Ruggero G. Pensa, Rosa Meo
IDA2
2009 Parameter-Free Hierarchical Co-clustering by n-Ary Splits
Dino Ienco, Ruggero G. Pensa, Rosa Meo
ECML/PKDD (1)2
2008 Constrained Co-clustering of Gene Expression Data
abstract
In many applications, the expert interpretation of co-clustering is easier than for mono-dimensional clustering. Co-clustering aims at computing a bi-partition that is a collection of co-clusters: each co-cluster is a group of objects associated to a group of attributes and these associations can support interpretations. Many constrained clustering algorithms have been proposed to exploit the domain knowledge and to improve partition relevancy in the mono-dimensional case (e.g., using the so-called must-link and cannot-link constraints). Here, we consider constrained co-clustering not only for extended must-link and cannot-link constraints (i.e., both objects and attributes can be involved), but also for interval constraints that enforce properties of co-clusters when considering ordered domains. We propose an iterative co-clustering algorithm which exploits user-defined constraints while minimizing the sum-squared residues, i.e., an objective function introduced for gene expression data clustering by Cho et al. (2004). We illustrate the added value of our approach in two applications on gene expression data.
Ruggero G. Pensa, Jean-François Boulicaut
SDM1
2008 SQUAT: A web tool to mine human, murine and avian SAGE data
abstract
BACKGROUND: There is an increasing need in transcriptome research for gene expression data and pattern warehouses. It is of importance to integrate in these warehouses both raw transcriptomic data, as well as some properties encoded in these data, like local patterns. DESCRIPTION: We have developed an application called SQUAT (SAGE Querying and Analysis Tools) which is available at: http://bsmc.insa-lyon.fr/squat/. This database gives access to both raw SAGE data and patterns mined from these data, for three species (human, mouse and chicken). This database allows to make simple queries like "In which biological situations is my favorite gene expressed?" as well as much more complex queries like: < >. Connections with external web databases enrich biological interpretations, and enable sophisticated queries. To illustrate the power of SQUAT, we show and analyze the results of three different queries, one of which led to a biological hypothesis that was experimentally validated. CONCLUSION: SQUAT is a user-friendly information retrieval platform, which aims at bringing some of the state-of-the-art mining tools to biologists.
Johan Leyritz, Stéphane Schicklin, Sylvain Blachon, Céline Keime, Céline Robardet, Jean-François Boulicaut, Jérémy Besson, Ruggero G. Pensa, Olivier Gandrillon
BMC Bioinform.8
2006 Towards Constrained Co-clustering in Ordered 0/1 Data Sets
Ruggero G. Pensa, Céline Robardet, Jean-François Boulicaut
ISMIS1
2006 Supporting bi-cluster interpretation in 0/1 data by means of local patterns
Ruggero G. Pensa, Céline Robardet, Jean-François Boulicaut
Intell. Data Anal.1
2005 From Local Pattern Mining to Relevant Bi-cluster Characterization
Ruggero G. Pensa, Jean-François Boulicaut
IDA1
2005 A Bi-clustering Framework for Categorical Data
Ruggero G. Pensa, Céline Robardet, Jean-François Boulicaut
PKDD1
2004 A Methodology for Biologically Relevant Pattern Discovery from Gene Expression Data
Ruggero G. Pensa, Jérémy Besson, Jean-François Boulicaut
Discovery Science1