Manuele Bicego

dblp:87/6321 · DBLP profile ↗
← Back
6ranked-venue papers in the field
4as first author
5since 2021 · last 2025
0000-0002-1008-3917ORCID · corroborated

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 5 (3 first)Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2025 Counterintuitive Behavior of Clustering Quality: Findings for K-Means on Synthetic and Real Data
Marco Loog, Jesse H. Krijthe, Manuele Bicego
IDA3
2025 TSRF-Dist: a novel time series distance based on extremely randomized canonical interval forests
abstract
Abstract This paper presents , a novel distance between time series based on Random Forests (RFs). We extend to the time-series domain concepts and tools of RF distances, a recent class of robust data-dependent distances defined for vectorial representations, thus proposing the first RF distance for time series. The distance is determined by (i) creating an RF to model a set of time series, and (ii) exploiting the trained RF to quantify the similarity between time series. As for the first step, we introduce in this paper the Extremely Randomized Canonical Interval Forest (ERCIF), a novel extension of Canonical Interval Forests that can model time series and can be trained without labels. We then exploit three different schemes, following ideas already employed in the vectorial case. The proposed distance, in different variants, has been thoroughly evaluated with 128 datasets from the archive, showing promising results compared with literature alternatives.
Alberto Azzari, Manuele Bicego, Carlo Combi, Andrea Cracco, Pietro Sala
Data Min. Knowl. Discov.2
2024 Computing Random Forest-distances in the presence of missing data
abstract
In this article, we study the problem of computing Random Forest-distances in the presence of missing data. We present a general framework which avoids pre-imputation and uses in an agnostic way the information contained in the input points. We centre our investigation on RatioRF, an RF-based distance recently introduced in the context of clustering and shown to outperform most known RF-based distance measures. We also show that the same framework can be applied to several other state-of-the-art RF-based measures and provide their extensions to the missing data case. We provide significant empirical evidence of the effectiveness of the proposed framework, showing extensive experiments with RatioRF on 15 datasets. Finally, we also positively compare our method with many alternative literature distances, which can be computed with missing values.
Manuele Bicego, Ferdinando Cicalese
ACM Trans. Knowl. Discov. Data1
2023 On the Good Behaviour of Extremely Randomized Trees in Random Forest-Distance Computation
Manuele Bicego, Ferdinando Cicalese
ECML/PKDD (4)1
2023 RatioRF: A Novel Measure for Random Forest Clustering Based on the Tversky's Ratio Model
abstract
In this paper we propose RatioRF, a novel Random Forest-based similarity measure for clustering. We build upon Tversky's ratio model definition of similarity and specialize it to the Random Forest case. We study some properties of the proposed axiomatic similarity measure and present an extensive experimental clustering analysis involving different datasets and configurations. Results confirm that RatioRF represents a good alternative to other similar measures for clustering recently studied in the literature.
Manuele Bicego, Ferdinando Cicalese, Antonella Mensi
IEEE Trans. Knowl. Data Eng.1
2020 Dissimilarity Random Forest Clustering
abstract
In this paper we present DisRFC (Dissimilarity Random Forest Clustering), a novel Random Forest Clustering approach which, contrarily to current methods which require in input a vectorial representation, works only with dissimilarities, thus being applicable also to all those problems where a vectorial representation is not available but a descriptive dissimilarity measure can be computed. In the DisRFC approach objects to be clustered are first modelled with a novel RF variant called Unsupervised Dissimilarity Random Forest (UD-RF), which functioning mechanisms are both unsupervised and based on dissimilarities. The trained UD-RF is then used to project objects in a binary vectorial space, where effective K-means procedures can be used to obtain the final clustering. In the paper we present different variants of DisRFC, thoroughly and positively evaluated using 10 different problems.
Manuele Bicego
ICDM1