VLDB 2026 Research / reviewers in the wild / expert
Manuele Bicego
dblp:87/6321
· DBLP profile ↗
78ranked-venue papers
36as first author
11since 2021 · last 2026
0000-0002-1008-3917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 30 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorSystems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Probabilistic Learnability of Compact Neural Network Preimage BoundsabstractAlthough recent provable methods have been developed to compute preimage bounds for neural networks, their scalability is fundamentally limited by the #P-hardness of the problem. In this work, we adopt a novel probabilistic perspective, aiming to deliver solutions with high-confidence guarantees and bounded error. To this end, we investigate the potential of bootstrap-based and randomized approaches that are capable of capturing complex patterns in high-dimensional spaces, including input regions where a given output property holds. In detail, we introduce Random Forest Property Verifier (RF-ProVe), a method that exploits an ensemble of randomized decision trees to generate candidate input regions satisfying a desired output property and refines them through active resampling. Our theoretical derivations offer formal statistical guarantees on region purity and global coverage, providing a practical, scalable solution for computing compact preimage approximations in cases where exact solvers fail to scale. Luca Marzari, Manuele Bicego, Ferdinando Cicalese, Alessandro Farinelli |
AAAI | 2 |
| 2025 | Counterintuitive Behavior of Clustering Quality: Findings for K-Means on Synthetic and Real Data
Marco Loog, Jesse H. Krijthe, Manuele Bicego |
IDA | 3 |
| 2025 | TSRF-Dist: a novel time series distance based on extremely randomized canonical interval forestsabstractAbstract This paper presents , a novel distance between time series based on Random Forests (RFs). We extend to the time-series domain concepts and tools of RF distances, a recent class of robust data-dependent distances defined for vectorial representations, thus proposing the first RF distance for time series. The distance is determined by (i) creating an RF to model a set of time series, and (ii) exploiting the trained RF to quantify the similarity between time series. As for the first step, we introduce in this paper the Extremely Randomized Canonical Interval Forest (ERCIF), a novel extension of Canonical Interval Forests that can model time series and can be trained without labels. We then exploit three different schemes, following ideas already employed in the vectorial case. The proposed distance, in different variants, has been thoroughly evaluated with 128 datasets from the archive, showing promising results compared with literature alternatives. Alberto Azzari, Manuele Bicego, Carlo Combi, Andrea Cracco, Pietro Sala |
Data Min. Knowl. Discov. | 2 |
| 2024 | An Extension of Random Forest-Clustering Schemes Which Works with Partition-Level Constraints
Manuele Bicego, Hafiz Ahmad Hassan |
ICPR (24) | 1 |
| 2024 | Computing Random Forest-distances in the presence of missing dataabstractIn this article, we study the problem of computing Random Forest-distances in the presence of missing data. We present a general framework which avoids pre-imputation and uses in an agnostic way the information contained in the input points. We centre our investigation on RatioRF, an RF-based distance recently introduced in the context of clustering and shown to outperform most known RF-based distance measures. We also show that the same framework can be applied to several other state-of-the-art RF-based measures and provide their extensions to the missing data case. We provide significant empirical evidence of the effectiveness of the proposed framework, showing extensive experiments with RatioRF on 15 datasets. Finally, we also positively compare our method with many alternative literature distances, which can be computed with missing values. Manuele Bicego, Ferdinando Cicalese |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | On the Good Behaviour of Extremely Randomized Trees in Random Forest-Distance Computation
Manuele Bicego, Ferdinando Cicalese |
ECML/PKDD (4) | 1 |
| 2023 | Also for k-means: more data does not imply better performanceabstractAbstract Arguably, a desirable feature of a learner is that its performance gets better with an increasing amount of training data, at least in expectation. This issue has received renewed attention in recent years and some curious and surprising findings have been reported on. In essence, these results show that more data does actually not necessarily lead to improved performance—worse even, performance can deteriorate. Clustering, however, has not been subjected to such kind of study up to now. This paper shows that k-means clustering, a ubiquitous technique in machine learning and data mining, suffers from the same lack of so-called monotonicity and can display deterioration in expected performance with increasing training set sizes. Our main, theoretical contributions prove that 1-means clustering is monotonic, while 2-means is not even weakly monotonic, i.e., the occurrence of nonmonotonic behavior persists indefinitely, beyond any training sample size. For larger k, the question remains open. Marco Loog, Jesse H. Krijthe, Manuele Bicego |
Mach. Learn. | 3 |
| 2023 | DisRFC: a dissimilarity-based Random Forest Clustering approach
Manuele Bicego |
Pattern Recognit. | 1 |
| 2023 | Detecting outliers from pairwise proximities: Proximity isolation forestsabstractBecause outliers are very different from the rest of the data, it is natural to represent outliers by their distances to other objects. Furthermore, there are many scenarios in which only pairwise distances are known, and feature-based outlier detection methods cannot directly be applied. Considering these observations, and given the success of Isolation Forests for (feature-based) outlier detection , we propose Proximity Isolation Forest, a proximity-based extension. The methodology only requires a set of pairwise distances to work, making it suitable for different types of data. Analogously to Isolation Forest, outliers are detected via their early isolation in the trees; to encode the isolation we design nine training strategies, both random and optimized. We thoroughly evaluate the proposed approach on fifteen datasets, successfully assessing its robustness and suitability for the task; additionally we compare favourably to alternative proximity-based methods. Antonella Mensi, David M. J. Tax, Manuele Bicego |
Pattern Recognit. | 3 |
| 2023 | RatioRF: A Novel Measure for Random Forest Clustering Based on the Tversky's Ratio ModelabstractIn this paper we propose RatioRF, a novel Random Forest-based similarity measure for clustering. We build upon Tversky's ratio model definition of similarity and specialize it to the Random Forest case. We study some properties of the proposed axiomatic similarity measure and present an extensive experimental clustering analysis involving different datasets and configurations. Results confirm that RatioRF represents a good alternative to other similar measures for clustering recently studied in the literature. Manuele Bicego, Ferdinando Cicalese, Antonella Mensi |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Enhanced anomaly scores for isolation forests
Antonella Mensi, Manuele Bicego |
Pattern Recognit. | 2 |
| 2020 | Dissimilarity Random Forest ClusteringabstractIn this paper we present DisRFC (Dissimilarity Random Forest Clustering), a novel Random Forest Clustering approach which, contrarily to current methods which require in input a vectorial representation, works only with dissimilarities, thus being applicable also to all those problems where a vectorial representation is not available but a descriptive dissimilarity measure can be computed. In the DisRFC approach objects to be clustered are first modelled with a novel RF variant called Unsupervised Dissimilarity Random Forest (UD-RF), which functioning mechanisms are both unsupervised and based on dissimilarities. The trained UD-RF is then used to project objects in a binary vectorial space, where effective K-means procedures can be used to obtain the final clustering. In the paper we present different variants of DisRFC, thoroughly and positively evaluated using 10 different problems. Manuele Bicego |
ICDM | 1 |
| 2020 | On learning Random Forests for Random Forest-clusteringabstractIn this paper we study the poorly investigated problem of learning Random Forests for distance-based Random Forest clustering. We studied both classic schemes as well as alternative approaches, novel in this context. In particular, we investigated the suitability of Gaussian Density Forests, Random Forests specifically designed for density estimation. Further, we introduce a novel variant of Random Forest, based on an effective non parametric by-pass estimator of the Rényi entropy, which can be useful when the parametric assumption is too strict. An empirical evaluation involving different datasets and different RF-clustering strategies confirms that the learning step is crucial for RF-clustering. We also present a set of practical guidelines useful to determine the most suitable variant of RF-clustering according to the problem under examination. Manuele Bicego, Francisco Escolano |
ICPR | 1 |
| 2020 | PowerHC: non linear normalization of distances for advanced nearest neighbor classificationabstractIn this paper we investigate the exploitation of non linear scaling of distances for advanced nearest neighbor classification. Starting from the recently found relation between the Hypersphere Classifier (HC) [1] and the Adaptive Nearest Neighbor rule (ANN) [2], here we propose PowerHC, an improved version of HC in which distances are normalized using a non linear mapping; non linear scaling of data, whose usefulness for feature spaces has been already assessed, has been hardly investigated for distances. A thorough experimental evaluation, involving 24 datasets and a challenging real world scenario of seismic signal classification, confirms the suitability of the proposed approach. Manuele Bicego, Mauricio Orozco-Alzate |
ICPR | 1 |
| 2020 | Proximity Isolation ForestsabstractIsolation Forests are a very successful approach for solving outlier detection tasks. Isolation Forests are based on classical Random Forest classifiers that require feature vectors as input. There are many situations where vectorial data is not readily available, for instance when dealing with input sequences or strings. In these situations, one can extract higher level characteristics from the input, which is typically hard and often loses valuable information. An alternative is to define a proximity between the input objects, which can be more intuitive. In this paper we propose the Proximity Isolation Forests that extend the Isolation Forests to non-vectorial data. The introduced methodology has been thoroughly evaluated on 8 different problems and it achieves very good results also when compared to other techniques. Antonella Mensi, Manuele Bicego, David M. J. Tax |
ICPR | 2 |
| 2020 | A cheaper Rectified-Nearest-Feature-Line-Segment classifier based on safe pointsabstractThe Rectified Nearest Feature Line Segment (RN-FLS) classifier is an improved version of the Nearest Feature Line (NFL) classification rule. RNFLS corrects two drawbacks of NFL, namely the interpolation and extrapolation inaccuracies, by applying two consecutive processes-segmentation and rectification - to the initial set of feature lines. The main drawbacks of this technique, occurring in both training and test phases, are the high computational cost of the rectification procedure and the exponential explosion of the number of lines. We propose a cheaper version of RNFLS, based on a characterization of the points that should form good lines. The characterization relies on a recent neighborhood-based principle that categorizes objects into four types: safe, borderline, rare and outliers, depending on the position of each point with respect to the other classes. The proposed approach represents a variant of RNFLS in the sense that it only considers lines between safe points. This allows a drastic reduction in the computational burden imposed by RNFLS. We carried out an empirical and thorough analysis based on different public data sets, showing that our proposed approach, in general, is not significantly different from RNFLS, but cheaper since the consideration of likely irrelevant feature line segments is avoided. Mauricio Orozco-Alzate, Manuele Bicego |
ICPR | 2 |
| 2020 | Time series segmentation for state-model generation of autonomous aquatic drones: A systematic framework
Alberto Castellini, Manuele Bicego, Francesco Masillo, Maddalena Zuccotto, Alessandro Farinelli |
Eng. Appl. Artif. Intell. | 2 |
| 2020 | Biclustering with dominant sets
Matteo Denitto, Manuele Bicego, Alessandro Farinelli, Sebastiano Vascon, Marcello Pelillo |
Pattern Recognit. | 2 |
| 2019 | K-Random Forests: a K-means style algorithm for Random Forest clusteringabstractIn this paper we present a novel clustering approach based on Random Forests, a popular classification and regression technique whose usability in the clustering scenario has been investigated to a lesser extent. In the clustering context, the most used class of approaches is based on the exploitation of a single Random Forest to derive a proximity measure between points, to be used with any distance-based clustering technique. On the contrary, our scheme exploits a set of Random Forests, each one devoted to model one cluster, in a spirit similar to that of the mixture models approach. These Random Forests, which provide flexible cluster descriptors, are iteratively updated using a K-means-like clustering algorithm. The proposed scheme, which we call K-Random Forests (K-RF), has been evaluated on five datasets: the obtained results suggest that it represents a valid alternative to classic Random Forest clustering algorithms as well as to other established clustering approaches. Manuele Bicego |
IJCNN | 1 |
| 2019 | Subspace Clustering for Situation Assessment in Aquatic Drones: A Sensitivity Analysis for State-Model ImprovementabstractIn this paper, we propose the use of subspace clustering to detect the states of dynamical systems from sequences of observations. In particular, we generate sparse and interpretable models that relate the states of aquatic drones involved in autonomous water monitoring to the properties (e.g., statistical distribution) of data collected by drone sensors. The subspace clustering algorithm used is called SubCMedians. A quantitative experimental analysis is performed to investigate the connections between i) learning parameters and performance, ii) noise in the data and performance. The clustering obtained with this analysis outperforms those generated by previous approaches. Alberto Castellini, Manuele Bicego, Domenico Daniele Bloisi, Jason Blum, Francesco Masillo, Sergio Peignier, Alessandro Farinelli |
Cybern. Syst. | 2 |
| 2019 | Orienteering-based informative path planning for environmental monitoring
Lorenzo Bottarelli, Manuele Bicego, Jason Blum, Alessandro Farinelli |
Eng. Appl. Artif. Intell. | 2 |
| 2019 | On the importance of local and global analysis in the judgment of similarity and dissimilarity of faces
Manuele Bicego, Enrico Grosso |
Image Vis. Comput. | 1 |
| 2019 | A dissimilarity-based multiple instance learning approach for protein remote homology detection
Antonella Mensi, Manuele Bicego, Pietro Lovato, Marco Loog, David M. J. Tax |
Pattern Recognit. Lett. | 2 |
| 2018 | Mining NMR Spectroscopy Using Topic ModelsabstractPattern Recognition techniques have been successfully exploited for the biomedical analysis of NMR spectra. In this context, it is crucial to derive a suitable representation for the data: among others, a successful line of research exploits the Bag of Words representation (called here “Bag of Peaks”). However, despite its success, the Bag of Peaks paradigm has not been fully explored: for example, appropriate probabilistic models (such as topic models) can further distill the information contained in the Bag of Words, allowing for more interpretable and accurate solutions for the task-at-hand. This paper is aimed at filling this gap, by investigating the usefulness of topic models in the analysis of NMR spectra. In particular, we first introduce an unsupervised approach, based on topic models, that performs soft biclustering of NMR spectra-this kind of unsupervised analysis being new in the NMR literature. Second, we show that descriptors extracted from topic models can be successfully employed for classification of NMR samples: compared to the original Bag of Words, we prove that our descriptors provide higher accuracies. Finally, we perform an empirical evaluation involving a complex dataset of spectra derived from fruits, and two datasets of medical NMR spectra: our analysis confirms the suitability of such models in the NMR spectra analysis. Manuele Bicego, Pietro Lovato, Marco DeBona, Flavia Guzzo, Michael Assfalg |
ICPR | 1 |
| 2018 | Clustering via binary embedding
Manuele Bicego, Mário A. T. Figueiredo |
Pattern Recognit. | 1 |
| 2018 | On the distinctiveness of the electricity load profile
Manuele Bicego, Alessandro Farinelli, Enrico Grosso, D. Paolini, Sarvapali D. Ramchurn |
Pattern Recognit. | 1 |
| 2018 | Biclustering with a quantum annealer
Lorenzo Bottarelli, Manuele Bicego, Matteo Denitto, Alessandra Di Pierro, Alessandro Farinelli, Riccardo Mengoni |
Soft Comput. | 2 |
| 2017 | Region-Based Correspondence Between 3D Shapes via Spatially Smooth Biclustering
Matteo Denitto, Simone Melzi, Manuele Bicego, Umberto Castellani, Alessandro Farinelli, Mário A. T. Figueiredo, Yanir Kleiman, Maks Ovsjanikov |
ICCV | 3 |
| 2017 | A hierarchical clustering approach to large-scale near-optimal coalition formation with quality guarantees
Alessandro Farinelli, Manuele Bicego, Filippo Bistaffa, Sarvapali D. Ramchurn |
Eng. Appl. Artif. Intell. | 2 |
| 2017 | Spike and slab biclustering
Matteo Denitto, Manuele Bicego, Alessandro Farinelli, Mário A. T. Figueiredo |
Pattern Recognit. | 2 |
| 2017 | A biclustering approach based on factor graphs and the max-sum algorithm
Matteo Denitto, Alessandro Farinelli, Mário A. T. Figueiredo, Manuele Bicego |
Pattern Recognit. | 4 |
| 2017 | Soft Ngram Representation and Modeling for Protein Remote Homology DetectionabstractRemote homology detection represents a central problem in bioinformatics, where the challenge is to detect functionally related proteins when their sequence similarity is low. Recent solutions employ representations derived from the sequence profile, obtained by replacing each amino acid of the sequence by the corresponding most probable amino acid in the profile. However, the information contained in the profile could be exploited more deeply, provided that there is a representation able to capture and properly model such crucial evolutionary information. In this paper, we propose a novel profile-based representation for sequences, called soft Ngram. This representation, which extends the traditional Ngram scheme (obtained by grouping N consecutive amino acids), permits considering all of the evolutionary information in the profile: this is achieved by extracting Ngrams from the whole profile, equipping them with a weight directly computed from the corresponding evolutionary frequencies. We illustrate two different approaches to model the proposed representation and to derive a feature vector, which can be effectively used for classification using a support vector machine (SVM). A thorough evaluation on three benchmarks demonstrates that the new approach outperforms other Ngram-based methods, and shows very promising results also in comparison with a broader spectrum of techniques. Pietro Lovato, Marco Cristani, Manuele Bicego |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2016 | Skeleton-Based Orienteering for Level Set EstimationabstractIn recent years, the use of unmanned vehicles for monitoring spatial environmental phenomena has gained increasing attention. Within this context, an interesting problem is level set estimation, where the goal is to identify regions of space where the analyzed phenomena (for example the PH value in a body of water) is above or below a given threshold level. Typically, in the literature this problem is approached with techniques which search for the most interesting sampling locations to collect the desired information (i.e., locations where we can gain the most information by sampling). However, the common assumption underlying this class of approaches is that acquiring a sample is expensive (e.g., in terms of consumed energy and time). In this paper, we take a different perspective on the same problem by considering the case where a mobile vehicle can continuously acquire measurements with a negligible cost, through high rate sampling sensors. In this scenario, it is crucial to reduce the path length that the mobile platform executes to collect the data. To address this issue, we propose a novel algorithm, called Skeleton-Based Orienteering for Level Set Estimation (SBOLSE). Our approach starts from the LSE formulation introduced in [10] and formulates the level set estimation problem as an orienteering problem. This allows one to determine informative locations while considering the length of the path. To combat the complexity associated with the orienteering approach, we propose a heuristic approach based on the concept of topological skeletonization. We evaluate our algorithm by comparing it with the state of the art approaches (i.e., LSE and LSE-batch) both on a real world dataset extracted from mobile platforms and on a synthetic dataset extracted from CO2 maps. Results show that our approach achieves a near optimal classification accuracy while significantly reducing the travel distance (up to 70% w.r.t LSE and 30% w.r.t. LSE-batch). Lorenzo Bottarelli, Manuele Bicego, Jason Blum, Alessandro Farinelli |
ECAI | 2 |
| 2016 | Weighted K-Nearest Neighbor revisitedabstractIn this paper we show that weighted K-Nearest Neighbor, a variation of the classic K-Nearest Neighbor, can be reinterpreted from a classifier combining perspective, specifically as a fixed combiner rule, the sum rule. Subsequently, we experimentally demonstrate that it can be rather beneficial to consider other combining schemes as well. In particular, we focus on trained combiners and illustrate the positive effect these can have on classification performance. Manuele Bicego, Marco Loog |
ICPR | 1 |
| 2016 | Traveling on discrete embeddings of gene expression
Pietro Lovato, Manuele Bicego, Maria Kesa, Nebojsa Jojic, Vittorio Murino, Alessandro Perina |
Artif. Intell. Medicine | 2 |
| 2016 | A bioinformatics approach to 2D shape classification
Manuele Bicego, Pietro Lovato |
Comput. Vis. Image Underst. | 1 |
| 2016 | Properties of the Box-Cox transformation for pattern classification
Manuele Bicego, Sisto Baldo |
Neurocomputing | 1 |
| 2015 | Biclustering Gene Expressions Using Factor Graphs and the Max-Sum Algorithm
Matteo Denitto, Alessandro Farinelli, Manuele Bicego |
IJCAI | 3 |
| 2015 | A Multimodal Approach for Protein Remote Homology DetectionabstractProtein remote homology detection represents a crucial and challenging task in bioinformatics: even if effective methods appeared in recent years, in several cases a proper characterization of remote evolutionary correlation can not be derived. In such situations, it may be possible that information derived from other sources helps, provided that it is possible to properly integrate such (even partial) information into existing models. In this paper, we provide some evidence that this route is feasible: inspired by the multimodal retrieval literature, we show how it is possible to exploit a simple multimodal approach to improve a model learned from a set of sequences, by using knowledge derived from a partial set of corresponding 3D structures. We investigate (with the SCOP 1.53 benchmark) the suitability of the proposed multimodal scheme, showing that a beneficial effect can be obtained even when a very reduced amount of structures are available. A further detailed analysis on a member of the GPCR superfamily confirms that this multimodal approach can extract information that cannot be obtained from sequence-based techniques. Pietro Lovato, Alejandro Giorgetti, Manuele Bicego |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | A Comparison between Time-Frequency and Cepstral Feature Representations for the Classification of Seismic-Volcanic Signals
Paola Castro-Cabrera, Mauricio Orozco-Alzate, Andrea Adami, Manuele Bicego, John Makario Londoño-Bonilla, Germán Castellanos-Domínguez |
CIARP | 4 |
| 2014 | Behavioural Biometrics Using Electricity Load ProfilesabstractModelling behavioural biometric patterns is a key issue for modern user centric applications, aimed at better monitoring users' activities, understanding their habits and detecting their identity. Following this trend, this paper investigates whether the electrical energy consumption of a user can be a distinctive behavioural biometric trait. In particular we analyse daily and weekly load profiles showing that they are closely related to the identity of the users. Hence, we believe that this level of analysis can open interesting application scenarios in the field of energy management and it provides a good working framework for the continuous development of smart environments with demonstrable benefits on real-world implementations. Manuele Bicego, F. Recchia, Alessandro Farinelli, Sarvapali D. Ramchurn, Enrico Grosso |
ICPR | 1 |
| 2014 | S-BLOSUM: Classification of 2D Shapes with Biological Sequence AlignmentabstractRecent works investigated the possibility to design solutions for pattern recognition problems by exploiting the huge amount of work done in bioinformatics. If the pattern recognition problem is cast in biological terms, then a huge range of algorithms, exploitable for classification, detection, visualization, etc. can be effectively borrowed. In this paper, we exploit biological sequence alignment tools to classify 2D shapes, tailoring the biological parameters of these tools to account for the different semantic of the 2D shape scenario. In particular, we propose a novel substitution matrix, which is the crucial parameter determining the sequence alignment solution. The new matrix, called S-BLOSUM, learns the rates of matches/mismatches in conserved portions of shapes belonging to the same category, and incorporates prior knowledge on the chosen representation for the 2D shape. On one hand, the experimental evaluation showed that the S-BLOSUM provides a significant improvement over the biological counterpart (BLOSUM), on the other hand, classification results prove that our approach is competitive with respect to the state of the art. Pietro Lovato, Alessio Milanese, Cesare Centomo, Alejandro Giorgetti, Manuele Bicego |
ICPR | 5 |
| 2014 | Expression Microarray Data Classification Using Counting Grids and Fisher KernelabstractHybrid generative-discriminative models are useful in biomedical applications- generative modeling extracts interpretable features from raw data, highlighting its properties and increasing classification accuracy when used as input for a discriminative classifier. This raises the question: which generative model should be used for a particular application? In this paper we apply a recently proposed generative model called the Counting Grid to expression microarray data and derive the corresponding Fisher kernel. We justify why this model is particularly well-suited for this application and evaluate classification accuracy on four gene expression data sets, including three tumor data sets and a blood sample data set from schizophrenic patients and healthy controls. We report state of the art results on three of the analyzed data sets and closely match the accuracy from previous work on the other. Alessandro Perina, Maria Kesa, Manuele Bicego |
ICPR | 3 |
| 2014 | Generative embeddings based on Rician mixtures for kernel-based classification of magnetic resonance images
Anna C. Carli, Mário A. T. Figueiredo, Manuele Bicego, Vittorio Murino |
Neurocomputing | 3 |
| 2014 | Faved! Biometrics: Tell Me Which Image You Like and I'll Tell You Who You AreabstractThis paper builds upon the belief that every human being has a built-in image aesthetic evaluation system. This sort of personal aesthetics mostly follows certain aesthetic rules widely studied in image aesthetics (e.g., rules of thirds, colorfulness, etc.), though it likely contains some innate, unique preferences. This paper is a proof of concept of this intuition, presenting personal aesthetics as a novel behavioral biometrical trait. In our scenario, personal aesthetics activate when an individual is presented with a set of photos he may like or dislike. The goal is to distill and encode the uniqueness of his visual preferences into a compact template. To this aim, we extract a pool of low- and high-level state-of-the-art image features from a set of Flickr images preferred by a user, feeding them successively into a LASSO regressor. LASSO highlights the most discriminant cues for the individual, allowing authentication and recognition tasks. The results are surprising given only 1 image as test. We can match the user identity against a gallery of 200 individuals definitely much better than chance. Using 20 images (all preferred by a single user) as a biometrical trait, we reach an AUC of 96%, considering the cumulative matching characteristic curve. Extensive experiments also support the interpretability of our approach, effectively modeling what is the “what we like” that distinguishes us from others. Pietro Lovato, Manuele Bicego, Cristina Segalin, Alessandro Perina, Nicu Sebe, Marco Cristani |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | C-Link: A Hierarchical Clustering Approach to Large-scale Near-optimal Coalition Formation
Alessandro Farinelli, Manuele Bicego, Sarvapali D. Ramchurn, Mauro Zucchelli |
IJCAI | 2 |
| 2013 | Documents as multiple overlapping windows into grids of countsabstractIn text analysis documents are represented as disorganized bags of words, models of count features are typically based on mixing a small number of topics \cite{lda,sam}. Recently, it has been observed that for many text corpora documents evolve into one another in a smooth way, with some features dropping and new ones being introduced. The counting grid \cite{cgUai} models this spatial metaphor literally: it is multidimensional grid of word distributions learned in such a way that a document's own distribution of features can be modeled as the sum of the histograms found in a window into the grid. The major drawback of this method is that it is essentially a mixture and all the content much be generated by a single contiguous area on the grid. This may be problematic especially for lower dimensional grids. In this paper, we overcome to this issue with the \emph{Componential Counting Grid} which brings the componential nature of topic models to the basic counting grid. We also introduce a generative kernel based on the document's grid usage and a visualization strategy useful for understanding large text corpora. We evaluate our approach on document classification and multimodal retrieval obtaining state of the art results on standard benchmarks. Alessandro Perina, Nebojsa Jojic, Manuele Bicego, Andrzej Truski |
NIPS | 3 |
| 2013 | Combining information theoretic kernels with generative embeddings for classification
Manuele Bicego, Aydin Ulas, Umberto Castellani, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
Neurocomputing | 1 |
| 2013 | Classification of Seismic Volcanic Signals Using Hidden-Markov-Model-Based Generative EmbeddingsabstractThe automated classification of seismic volcanic signals has been faced with several different pattern recognition approaches. Among them, hidden Markov models (HMMs) have been advocated as a cost-effective option having the advantages of a straightforward Bayesian interpretation and the capacity of dealing with seismic sequences of different lengths. In the volcano seismology scenario, HMM-based classification schemes were only based on a standard and purely generative scheme, i.e., the Bayes rule: training an HMM per class and classifying an incoming seismic signal according to the class whose model shows the highest likelihood. In this paper, a novel HMM-based classification approach for pretriggered seismic volcanic signals is proposed. The main idea is to enrich the classical HMM scheme with a discriminative step that is able to recover from situations when the classical Bayes classification rule is not sufficient. More in detail, a generative embedding scheme is used, which employs the models to map the signals into a vector space, which is called generative embedding space. In such a space, any discriminative vector-based classifier can be applied. A thorough set of experiments, which is carried out on pretriggered signals recorded at Galeras Volcano in Colombia, shows that the proposed approach typically outperforms standard HMM-based classification schemes, also in some cross-station cases. Manuele Bicego, Carolina Acosta-Munoz, Mauricio Orozco-Alzate |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | Tell Me What You Like and I'll Tell You What You Are: Discriminating Visual Preferences on Flickr Data
Pietro Lovato, Alessandro Perina, Nicu Sebe, Omar Zandonà, Alessio Montagnini, Manuele Bicego, Marco Cristani |
ACCV (1) | 6 |
| 2012 | Automatic Classification of Volcanic Earthquakes in HMM-Induced Vector Spaces
Riccardo Avesani, Alessio Azzoni, Manuele Bicego, Mauricio Orozco-Alzate |
CIARP | 3 |
| 2012 | 2D shape recognition using biological sequence alignment tools
Manuele Bicego, Pietro Lovato |
ICPR | 1 |
| 2012 | Generative Embeddings based on Rician Mixtures - Application to Kernel-based Discriminative Classification of Magnetic Resonance Images
Anna C. Carli, Mário A. T. Figueiredo, Manuele Bicego, Vittorio Murino |
ICPRAM (1) | 3 |
| 2012 | Investigating Topic Models' Capabilities in Expression Microarray Data ClassificationabstractIn recent years a particular class of probabilistic graphical models-called topic models-has proven to represent an useful and interpretable tool for understanding and mining microarray data. In this context, such models have been almost only applied in the clustering scenario, whereas the classification task has been disregarded by researchers. In this paper, we thoroughly investigate the use of topic models for classification of microarray data, starting from ideas proposed in other fields (e.g., computer vision). A classification scheme is proposed, based on highly interpretable features extracted from topic models, resulting in a hybrid generative-discriminative approach; an extensive experimental evaluation, involving 10 different literature benchmarks, confirms the suitability of the topic models for classifying expression microarray data. Manuele Bicego, Pietro Lovato, Alessandro Perina, Marianna Fasoli, Massimo Delledonne, Mario Pezzotti, Annalisa Polverari, Vittorio Murino |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2011 | Multimodal Schizophrenia Detection by Multiclassification Analysis
Aydin Ulas, Umberto Castellani, Pasquale Mirtuono, Manuele Bicego, Vittorio Murino, Stefania Cerruti, Marcella Bellani, Manfredo Atzori, Gianluca Rambaldelli, Michele Tansella, Paolo Brambilla |
CIARP | 4 |
| 2010 | Combining free energy score spaces with information theoretic kernels: Application to scene classificationabstractMost approaches to learn classifiers for structured objects (e.g., images) use generative models in a classical Bayesian framework. However, state-of-the-art classifiers for vectorial data (e.g., support vector machines) are learned discriminatively. A generative embedding is a mapping from the object space into a fixed dimensional score space, induced by a generative model, usually learned from data. The fixed dimensionality of these generative score spaces makes them adequate for discriminative learning of classifiers, thus bringing together the best of the discriminative and generative paradigms. In particular, it was recently shown that this hybrid approach outperforms a classifier obtained directly for the generative model upon which the score space was built. Using a generative embedding involves two steps: (i) defining and learning the generative model and using it to build the embedding; (ii) discriminatively learning a (maybe kernel) classifier on the adopted score space. The literature on generative embeddings is essentially focused on step (i), usually using some standard off-the-shelf tool for step (ii). In this paper, we adopt a different approach, by focusing also on the discriminative learning step. In particular, we combine two very recent and top performing tools in each of the steps: (i) the free energy score space; (ii) non-extensive information theoretic kernels. In this paper, we apply this methodology in scene recognition. Experimental results on two benchmark datasets shows that our approach yields state-of-the-art performance. Manuele Bicego, Alessandro Perina, Vittorio Murino, André F. T. Martins, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
ICIP | 1 |
| 2010 | Biclustering of Expression Microarray Data with Topic ModelsabstractThis paper presents an approach to extract biclusters from expression micro array data using topic models - a class of probabilistic models which allow to detect interpretable groups of highly correlated genes and samples. Starting from a topic model learned from the expression matrix, some automatic rules to extract biclusters are presented, which overcome the drawbacks of previous approaches. The methodology has been positively tested with synthetic benchmarks, as well as with a real experiment involving two different species of grape plants (Vitis vinifera and Vitis riparia). Manuele Bicego, Pietro Lovato, Alberto Ferrarini, Massimo Delledonne |
ICPR | 1 |
| 2010 | 2D Shape Recognition Using Information Theoretic KernelsabstractIn this paper, a novel approach for contour based 2D shape recognition is proposed, using a class of information theoretic kernels recently introduced. This kind of kernels, based on a non-extensive generalization of the classical Shannon information theory, are defined on probability measures. In the proposed approach, chain code representations are first extracted from the contours; then n-gram statistics are computed and used as input to the information theoretic kernels. We tested different versions of such kernels, using support vector machine and nearest neighbor classifiers. An experimental evaluation on the Chicken pieces dataset shows that the proposed approach significantly outperforms the current state-of-the-art methods. Manuele Bicego, André F. T. Martins, Vittorio Murino, Pedro M. Q. Aguiar, Mário A. T. Figueiredo |
ICPR | 1 |
| 2010 | Nonlinear Mappings for Generative Kernels on Latent Variable ModelsabstractGenerative kernels have emerged in the last years as an effective method for mixing discriminative and generative approaches. In particular, in this paper, we focus on kernels defined on generative models with latent variables (e.g. the states in a Hidden Markov Model). The basic idea underlying these kernels is to compare objects, via a inner product, in a feature space where the dimensions are related to the latent variables of the model. Here we propose to enhance these kernels via a nonlinear normalization of the space, namely a nonlinear mapping of space dimensions able to exploit their discriminative characteristics. In this paper we investigate three possible nonlinear mappings, for two HMM-based generative kernels, testing them in different sequence classification problems, with really promising results. Anna C. Carli, Manuele Bicego, Sisto Baldo, Vittorio Murino |
ICPR | 2 |
| 2009 | Online subjective feature selection for occlusion management in tracking applicationsabstractMost of the state-of-the-art tracking algorithms are prone to error when dealing with occlusions, especially when the involved moving objects are hardly discernible in appearance. In this paper, we propose a multi-object particle filtering tracking framework particularly suited to manage the occlusion problem. The presented solution consists in the introduction of a online subjective feature selection mechanism, which highlights and employs the most discriminant features characterizing a single object with respect to the neighbouring objects. The policy adopted fits formally in the observation step of the particle filtering process, it is effective and not computationally costly. Trials carried out on illustrative synthetic data and on recent challenging benchmark sequences report compelling performances and encourage further development of the technique. Loris Bazzani, Marco Cristani, Manuele Bicego, Vittorio Murino |
ICIP | 3 |
| 2009 | Bag of Peaks: interpretation of NMR spectrometryabstractMOTIVATION: The analysis of high-resolution proton nuclear magnetic resonance (NMR) spectrometry can assist human experts to implicate metabolites expressed by diseased biofluids. Here, we explore an intermediate representation, between spectral trace and classifier, able to furnish a communicative interface between expert and machine. This representation permits equivalent, or better, classification accuracies than either principal component analysis (PCA) or multi-dimensional scaling (MDS). In the training phase, the peaks in each trace are detected and clustered in order to compile a common dictionary, which could be visualized and adjusted by an expert. The dictionary is used to characterize each trace with a fixed-length feature vector, termed Bag of Peaks, ready to be classified with classical supervised methods. RESULTS: Our small-scale study, concerning Type I diabetes in Sardinian children, provides a preliminary indication of the effectiveness of the Bag of Peaks approach over standard PCA and MDS. Consistently, higher classification accuracies are obtained once a sufficient number of peaks (>10) are included in the dictionary. A large-scale simulation of noisy spectra further confirms this advantage. Finally, suggestions for metabolite-peak loci that may be implicated in the disease are obtained by applying standard feature selection techniques. Gavin Brelstaff, Manuele Bicego, Nicola Culeddu, Matilde Chessa |
Bioinform. | 2 |
| 2009 | Dynamic face recognition: From human to machine vision
Massimo Tistarelli, Manuele Bicego, Enrico Grosso |
Image Vis. Comput. | 2 |
| 2009 | Soft clustering using weighted one-class support vector machines
Manuele Bicego, Mário A. T. Figueiredo |
Pattern Recognit. | 1 |
| 2009 | Component-based discriminative classification for hidden Markov models
Manuele Bicego, Elzbieta Pekalska, David M. J. Tax, Robert P. W. Duin |
Pattern Recognit. | 1 |
| 2008 | Generalized Gaussian distributions for sequential data classificationabstractIt has been shown in many different contexts that the Generalized Gaussian (GG) distribution represents a flexible and suitable tool for data modeling. Almost all the reported applications are focused on modeling points (fixed length vectors); a different but crucial scenario, where the employment of the GG has received little attention, is the modeling of sequential data, i.e. variable length vectors. This paper explores this last direction, describing a variant of the well known Hidden Markov Model (HMM) where the emission probability function of each state is represented by a GG. A training strategy based on the Expectation Maximization (EM) algorithm is presented. Different experiments using both synthetic and real data (EEG signal classification and face recognition) show the suitability of the proposed approach compared with the standard Gaussian HMM. Manuele Bicego, Daniel González-Jiménez, Enrico Grosso, José Luis Alba-Castro |
ICPR | 1 |
| 2008 | Distinctiveness of faces: A computational approachabstractThis paper develops and demonstrates an original approach to face-image analysis based on identifying distinctive areas of each individual's face by its comparison to others in the population. The method differs from most others—that we refer as unary —where salient regions are defined by analyzing only images of the same individual. We extract a set of multiscale patches from each face image before projecting them into a common feature space. The degree of “distinctiveness” of any patch depends on its distance in feature space from patches mapped from other individuals. First a pairwise analysis is developed and then a simple generalization to the multiple-face case is proposed. A perceptual experiment, involving 45 observers, indicates the method to be fairly compatible with how humans mark faces as distinct. A quantitative example of face authentication is also performed in order to show the essential role played by the distinctive information. A comparative analysis shows that performance of our n-ary approach is as good as several contemporary unary, or binary, methods, while tapping a complementary source of information. Furthermore, we show it can also provide a useful degree of illumination invariance. Manuele Bicego, Enrico Grosso, Andrea Lagorio, Gavin Brelstaff, Linda Brodo, Massimo Tistarelli |
ACM Trans. Appl. Percept. | 1 |
| 2007 | Audio-Visual Event Recognition in Surveillance Video SequencesabstractIn the context of the automated surveillance field, automatic scene analysis and understanding systems typically consider only visual information, whereas other modalities, such as audio, are typically disregarded. This paper presents a new method able to integrate audio and visual information for scene analysis in a typical surveillance scenario, using only one camera and one monaural microphone. Visual information is analyzed by a standard visual background/foreground (BG/FG) modelling module, enhanced with a novelty detection stage and coupled with an audio BG/FG modelling scheme. These processes permit one to detect separate audio and visual patterns representing unusual unimodal events in a scene. The integration of audio and visual data is subsequently performed by exploiting the concept of synchrony between such events. The audio-visual (AV) association is carried out online and without need for training sequences, and is actually based on the computation of a characteristic feature called audio-video concurrence matrix, allowing one to detect and segment AV events, as well as to discriminate between them. Experimental tests involving classification and clustering of events show all the potentialities of the proposed approach, also in comparison with the results obtained by employing the single modalities and without considering the synchrony issue Marco Cristani, Manuele Bicego, Vittorio Murino |
IEEE Trans. Multim. | 2 |
| 2006 | Recognizing People's Faces: from Human to Machine VisionabstractAs confirmed by recent neurophysiological studies, the use of dynamic information is extremely important for humans in visual perception of biological forms and motion. Apart from the mere computation of the visual motion of the viewed objects, the motion itself conveys far more information, which helps understanding the scene. This paper provides an overview and some new insights on the use of dynamic visual information for face recognition. In this context, not only physical features emerge in the face representation, but also behavioral features should be accounted. While physical features are obtained from the subject's face appearance, behavioral features are obtained from the individual motion and articulation of the face. In order to capture both the face appearance and the face dynamics, a dynamical face model based on a combination of hidden Markov models is presented. The number of states (or facial expressions) are automatically determined from the data by unsupervised clustering of expressions of faces in the video. The underlying architecture closely recalls the neural patterns activated in the perception of moving faces. Preliminary results on real video image data show the feasibility of the proposed approach Massimo Tistarelli, Manuele Bicego, Enrico Grosso |
ICARCV | 2 |
| 2006 | Unsupervised scene analysis: A hidden Markov model approach
Manuele Bicego, Marco Cristani, Vittorio Murino |
Comput. Vis. Image Underst. | 1 |
| 2006 | Similarity-based pattern recognition
Manuele Bicego, Vittorio Murino, Marcello Pelillo, Andrea Torsello |
Pattern Recognit. | 1 |
| 2005 | A supervised data-driven approach for microarray spot quality classification
Manuele Bicego, Maria Rosario Martinez, Vittorio Murino |
Pattern Anal. Appl. | 1 |
| 2005 | A Hidden Markov Model approach for appearance-based 3D object recognition
Manuele Bicego, Umberto Castellani, Vittorio Murino |
Pattern Recognit. Lett. | 1 |
| 2004 | Audio-Video Integration for Background Modelling
Marco Cristani, Manuele Bicego, Vittorio Murino |
ECCV (2) | 2 |
| 2004 | Hybrid HMM/SVM Model for the Analysis and Segmentation of Teleoperation TasksabstractThe automatic execution of a complex task requires the identification of an underlying mental model to derive a possible task control sequence. The model aims at analysing and segmenting the task in simpler sub-tasks. As an example of a complex task, in this paper we consider teleoperation where a person commands a remote robot. This paper presents a new modeling approach using hidden Markov models (HMM) and support vector machines (SVM) to analyse the force/torque signals of a teleoperation task. The task is divided into simpler sub-tasks and the model is used to segment the signals in each sub-task. The segmentation gives informations on the system behavior identifying the changes of the model states. Peg in hole force/torque data are used for testing the model. The results are consistent with the literature with respect to off-line analysis, whereas a significant increase of performance is achieved for on-line analysis. Andrea Castellani, Debora Botturi, Manuele Bicego, Paolo Fiorini |
ICRA | 3 |
| 2004 | Investigating Hidden Markov Models' Capabilities in 2D Shape ClassificationabstractIn this paper, Hidden Markov Models (HMMs) are investigated for the purpose of classifying planar shapes represented by their curvature coefficients. In the training phase, special attention is devoted to the initialization and model selection issues, which make the learning phase particularly effective. The results of tests on different data sets show that the proposed system is able to accurately classify objects that were translated, rotated, occluded, or deformed by shearing, also in the presence of noise. Manuele Bicego, Vittorio Murino |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Similarity-based classification of sequences using hidden Markov models
Manuele Bicego, Vittorio Murino, Mário A. T. Figueiredo |
Pattern Recognit. | 1 |
| 2003 | Automatic road extraction from aerial images by probabilistic contour trackingabstractIn this paper a new automatic approach to road extraction from aerial images is proposed. This method improves a recently introduced promising approach to probabilistic contour tracking, originally semi-automatic, by adding a fully automatic initialization strategy and a merging methodology, able to combine the different obtained results. The initialization strategy is based on the Hough transform and on some topological considerations, and the merging step is based on a new introduced quality measure, based on color and gradient information. Experimental results on real highly complex images show that the proposed approach is a promising and fully automatic method for extracting roads from images, even in presence of highly urbanized areas, occlusions or shadows. Manuele Bicego, Silvio Dalfini, Gianni Vernazza, Vittorio Murino |
ICIP (3) | 1 |
| 2003 | A sequential pruning strategy for the selection of the number of states in hidden Markov models
Manuele Bicego, Vittorio Murino, Mário A. T. Figueiredo |
Pattern Recognit. Lett. | 1 |