Emanuele Sansone

dblp:138/0860 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 38% Representation and self-supervised learning · 21% Generative modeling · 13%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning
feature decorrelation
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
non-contrastive self-supervised learning
0.912025
Collapse-Proof Non-Contrastive Self-Supervised Learning · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › exponential family
categorical distribution
0.712023
Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model
0.712023
Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick · NeurIPS 2023
Machine learning › Optimization for machine learning
gradient estimation
0.712023
Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick · NeurIPS 2023
Machine learning › Optimization for machine learning
gradient estimator
0.712023
Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick · NeurIPS 2023
Machine learning › Reinforcement learning › policy optimization › policy gradient
REINFORCE
0.712023
Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
adaptive MCMC
0.612022
LSB: Local Self-Balancing MCMC in Discrete Spaces · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
discrete sampling
0.612022
LSB: Local Self-Balancing MCMC in Discrete Spaces · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
locally balanced proposals
0.612022
LSB: Local Self-Balancing MCMC in Discrete Spaces · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.612022
LSB: Local Self-Balancing MCMC in Discrete Spaces · ICML 2022
Machine learning › Representation and self-supervised learning › structured representation
neuro-symbolic representation
0.612022
VAEL: Bridging Variational Autoencoders and Probabilistic Logic Programming · NeurIPS 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic programming
probabilistic logic programming
0.612022
VAEL: Bridging Variational Autoencoders and Probabilistic Logic Programming · NeurIPS 2022
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.612022
LSB: Local Self-Balancing MCMC in Discrete Spaces · ICML 2022
Machine learning › Generative modeling
variational autoencoder
0.612022
VAEL: Bridging Variational Autoencoders and Probabilistic Logic Programming · NeurIPS 2022
Machine learning › Learning paradigms › weakly supervised learning
positive-unlabeled learning
0.412019
Efficient Training for Positive Unlabeled Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Image and video processing › image sequence processing
temporal alignment
0.312017
Automatic Synchronization of Multi-user Photo Galleries · IEEE Trans. Multim. 2017
Machine learning › Efficient and distributed learning
scalable learning
0.112019
Efficient Training for Positive Unlabeled Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2019

Methods — techniques the papers use, named apart from their topics

hyperdimensional computing · 0.9contrastive learning theory · 0.9log-derivative trick · 0.7catlog-derivative trick · 0.7variational inference · 0.6self-balancing learning · 0.6probabilistic logic programming · 0.6mutual information · 0.6end-to-end differentiable learning · 0.6optimization · 0.4probabilistic graphical model · 0.3minimum spanning tree · 0.3convolutional neural network features · 0.3
YearPublicationVenuePosition
2025 Collapse-Proof Non-Contrastive Self-Supervised Learning
abstract
We present a principled and simplified design of the projector and loss function for non-contrastive self-supervised learning based on hyperdimensional computing. We theoretically demonstrate that this design introduces an inductive bias that encourages representations to be simultaneously decorrelated and clustered, without explicitly enforcing these properties. This bias provably enhances generalization and suffices to avoid known training failure modes, such as representation, dimensional, cluster, and intracluster collapses. We validate our theoretical findings on image datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet-100. Our approach effectively combines the strengths of feature decorrelation and cluster-based self-supervised learning methods, overcoming training failure modes while achieving strong generalization in clustering and linear classification tasks.
Emanuele Sansone, Tim Lebailly, Tinne Tuytelaars
ICML1
2024 EXPLAIN, AGREE, LEARN: Scaling Learning for Neural Probabilistic Logic
abstract
Neural probabilistic logic systems follow the neuro-symbolic (NeSy) paradigm by combining the perceptive and learning capabilities of neural networks with the robustness of probabilistic logic. Learning corresponds to likelihood optimization of the neural networks. However, to obtain the likelihood exactly, expensive probabilistic logic inference is required. To scale learning to more complex systems, we therefore propose to instead optimize a sampling based objective. We prove that the objective has a bounded error with respect to the likelihood, which vanishes when increasing the sample count. Furthermore, the error vanishes faster by exploiting a new concept of sample diversity. We then develop the EXPLAIN, AGREE, LEARN (EXAL) method that uses this objective. EXPLAIN samples explanations for the data. AGREE reweighs each explanation in concordance with the neural component. LEARN uses the reweighed explanations as a signal for learning. In contrast to previous NeSy methods, EXAL can scale to larger problem sizes while retaining theoretical guarantees on the error. Experimentally, our theoretical claims are verified and EXAL outperforms recent NeSy methods when scaling up the MNIST addition and Warcraft pathfinding problems.
Victor Verreet, Lennert De Smet, Luc De Raedt, Emanuele Sansone
ECAI4
2023 Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative Trick
abstract
Categorical random variables can faithfully represent the discrete and uncertain aspects of data as part of a discrete latent variable model. Learning in such models necessitates taking gradients with respect to the parameters of the categorical probability distributions, which is often intractable due to their combinatorial nature. A popular technique to estimate these otherwise intractable gradients is the Log-Derivative trick. This trick forms the basis of the well-known REINFORCE gradient estimator and its many extensions. While the Log-Derivative trick allows us to differentiate through samples drawn from categorical distributions, it does not take into account the discrete nature of the distribution itself. Our first contribution addresses this shortcoming by introducing the CatLog-Derivative trick -- a variation of the Log-Derivative trick tailored towards categorical distributions. Secondly, we use the CatLog-Derivative trick to introduce IndeCateR, a novel and unbiased gradient estimator for the important case of products of independent categorical distributions with provably lower variance than REINFORCE. Thirdly, we empirically show that IndeCateR can be efficiently implemented and that its gradient estimates have significantly lower bias and variance for the same number of samples compared to the state of the art.
Lennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos Martires
NeurIPS2
2022 LSB: Local Self-Balancing MCMC in Discrete Spaces
abstract
We present the Local Self-Balancing sampler (LSB), a local Markov Chain Monte Carlo (MCMC) method for sampling in purely discrete domains, which is able to autonomously adapt to the target distribution and to reduce the number of target evaluations required to converge. LSB is based on (i) a parametrization of locally balanced proposals, (ii) an objective function based on mutual information and (iii) a self-balancing learning procedure, which minimises the proposed objective to update the proposal parameters. Experiments on energy-based models and Markov networks show that LSB converges using a smaller number of queries to the oracle distribution compared to recent local MCMC samplers.
Emanuele Sansone
ICML1
2022 VAEL: Bridging Variational Autoencoders and Probabilistic Logic Programming
abstract
We present VAEL, a neuro-symbolic generative model integrating variational autoencoders (VAE) with the reasoning capabilities of probabilistic logic (L) programming. Besides standard latent subsymbolic variables, our model exploits a probabilistic logic program to define a further structured representation, which is used for logical reasoning. The entire process is end-to-end differentiable. Once trained, VAEL can solve new unseen generation tasks by (i) leveraging the previously acquired knowledge encoded in the neural component and (ii) exploiting new logical programs on the structured latent space. Our experiments provide support on the benefits of this neuro-symbolic integration both in terms of task generalization and data efficiency. To the best of our knowledge, this work is the first to propose a general-purpose end-to-end framework integrating probabilistic logic programming into a deep generative model.
Eleonora Misino, Giuseppe Marra, Emanuele Sansone
NeurIPS3
2020 Coulomb Autoencoders
Emanuele Sansone, Hafiz Tiomoko Ali
ECAI1
2019 Efficient Training for Positive Unlabeled Learning
abstract
Positive unlabeled (PU) learning is useful in various practical situations, where there is a need to learn a classifier for a class of interest from an unlabeled data set, which may contain anomalies as well as samples from unknown classes. The learning task can be formulated as an optimization problem under the framework of statistical learning theory. Recent studies have theoretically analyzed its properties and generalization performance, nevertheless, little effort has been made to consider the problem of scalability, especially when large sets of unlabeled data are available. In this work we propose a novel scalable PU learning algorithm that is theoretically proven to provide the optimal solution, while showing superior computational and memory performance. Experimental evaluation confirms the theoretical evidence and shows that the proposed method can be successfully applied to a large variety of real-world problems involving PU learning.
Emanuele Sansone, Francesco G. B. De Natale, Zhi-Hua Zhou
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 Automatic Synchronization of Multi-user Photo Galleries
abstract
In this paper we address the issue of photo galleries synchronization, where pictures related to the same event are collected by different users. Existing solutions to address the problem are usually based on unrealistic assumptions, like time consistency across photo galleries, and often heavily rely on heuristics, therefore limiting the applicability to real-world scenarios. We propose a solution that achieves better generalization performance for the synchronization task compared to the available literature. The method is characterized by three stages: at first, deep convolutional neural network features are used to assess the visual similarity among the photos; then, pairs of similar photos are detected across different galleries and used to construct a graph; eventually, a probabilistic graphical model is used to estimate the temporal offset of each pair of galleries, by traversing the minimum spanning tree extracted from this graph. The experimental evaluation is conducted on four publicly available datasets covering different types of events, demonstrating the strength of our proposed method. A thorough discussion of the obtained results is provided for a critical assessment of the quality in synchronization.
Emanuele Sansone, Konstantinos Apostolidis, Nicola Conci, Giulia Boato, Vasileios Mezaris, Francesco G. B. De Natale
IEEE Trans. Multim.1
2016 Classtering: Joint Classification and Clustering with Mixture of Factor Analysers
abstract
In this work we propose a novel parametric Bayesian model for the problem of semi-supervised classification and clustering. Standard approaches of semi-supervised classification can recognize classes but cannot find groups of data. On the other hand, semi-supervised clustering techniques are able to discover groups of data but cannot find the associations between clusters and classes. The proposed model can classify and cluster samples simultaneously, allowing the analysis of data in the presence of an unknown number of classes and/or an arbitrary number of clusters per class. Experiments on synthetic and real world data show that the proposed model compares favourably to state-of-the-art approaches for semi-supervised clustering and that the discovered clusters can help to enhance classification performance, even in cases where the cluster and the low density separation assumptions do not hold. We finally show that when applied to a challenging real-world problem of subgroup discovery in breast cancer, the method is capable of maximally exploiting the limited information available and identifying highly promising subgroups.
Emanuele Sansone, Andrea Passerini, Francesco G. B. De Natale
ECAI1