Alessio Benavoli

dblp:12/2959 · DBLP profile ↗
← Back
41ranked-venue papers
27as first author
12since 2021 · last 2026
0000-0002-2522-7178ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 18 first-author · 11 since 2021Databases, data management, data science and information retrieval · 16 · 10 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorTheory of computation · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Connecting classical finite exchangeability to quantum theory and indistinguishability
Alessio Benavoli, Alessandro Facchini, Marco Zaffalon
Int. J. Approx. Reason.1
2024 Credal Valuation Network for Ongoing Threat Assessment
abstract
The paper develops a valuation network for sequential assessment of threat under epistemic uncertainty based on theoretical foundations and semantics of imprecise probability theory. The valuations are expressed as credal sets defined by coherent probability intervals on singletons. The combination rule is the generalized Bayes rule introduced by Walley. The model of a single-target threat is based on the classical “capability-intent” paradigm in an air surveillance context. Numerical results illustrate the performance of developed credal valuation network (with imprecise probabilities) against the valuation network with precise probabilistic models.
Branko Ristic 0001, Alessio Benavoli
FUSION2
2024 Linearly Constrained Gaussian Processes are SkewGPs: application to Monotonic Preference Learning and Desirability
abstract
We show that existing approaches to Linearly Constrained Gaussian Processes (LCGP) for regression, based on imposing constraints on a finite set of operational points, can be seen as Skew Gaussian Processes (SkewGPs). In particular, focusing on inequality constraints and building upon a recent unification of regression, classification, and preference learning through SkewGPs, we extend LCGP to handle monotonic preference learning and desirability, crucial for understanding and predicting human decision making. We demonstrate the efficacy of the proposed model on simulated and real data.
Alessio Benavoli, Dario Azzimonti
UAI1
2023 Learning Choice Functions with Gaussian Processes
abstract
In consumer theory, ranking available objects by means of preference relations yields the most common description of individual choices. However, preference-based models assume that individuals: (1) give their preferences only between pairs of objects; (2) are always able to pick the best preferred object. In many situations, they may be instead choosing out of a set with more than two elements and, because of lack of information and/or incomparability (objects with contradictory characteristics), they may not be able to select a single most preferred object. To address these situations, we need a choice model which allows an individual to express a set-valued choice. Choice functions provide such a mathematical framework. We propose a Gaussian Process model to learn choice functions from choice data. The model assumes a multiple utility representation of a choice function based on the concept of Pareto rationalization, and derives a strategy to learn both the number and the values of these latent multiple utilities. Simulation experiments demonstrate that the proposed model outperforms the state-of-the-art methods.
Alessio Benavoli, Dario Azzimonti, Dario Piga
UAI1
2023 Nonlinear desirability as a linear classification problem
abstract
This paper presents an interpretation as classification problem for standard desirability and other instances of nonlinear desirability (convex coherence and positive additive coherence). In particular, we analyze different sets of rationality axioms and, for each one of them, we show that proving that a subject respects these axioms on the basis of a finite set of acceptable and a finite set of rejectable gambles can be reformulated as a binary classification problem where the family of classifiers used changes with the axioms considered. Moreover, by borrowing ideas from machine learning, we show the possibility of defining a feature mapping, which allows us to reformulate the above nonlinear classification problems as linear ones in higher-dimensional spaces. This allows us to interpret gambles directly as payoffs vectors of monetary lotteries, as well as to provide a practical tool to check the rationality of an agent.
Arianna Casanova, Alessio Benavoli, Marco Zaffalon
Int. J. Approx. Reason.2
2023 Correlated product of experts for sparse Gaussian process regression
abstract
Gaussian processes (GPs) are an important tool in machine learning and statistics. However, off-the-shelf GP inference procedures are limited to datasets with several thousand data points because of their cubic computational complexity. For this reason, many sparse GPs techniques have been developed over the past years. In this paper, we focus on GP regression tasks and propose a new approach based on aggregating predictions from several local and correlated experts. Thereby, the degree of correlation between the experts can vary between independent up to fully correlated experts. The individual predictions of the experts are aggregated taking into account their correlation resulting in consistent uncertainty estimates. Our method recovers independent Product of Experts, sparse GP and full GP in the limiting cases. The presented framework can deal with a general kernel function and multiple variables, and has a time and space complexity which is linear in the number of experts and data samples, which makes our approach highly scalable. We demonstrate superior performance, in a time vs. accuracy sense, of our proposed method against state-of-the-art GP approximations for synthetic as well as several real-world datasets with deterministic and stochastic optimization. Supplementary Information: The online version contains supplementary material available at 10.1007/s10994-022-06297-3.
Manuel Schürch, Dario Azzimonti, Alessio Benavoli, Marco Zaffalon
Mach. Learn.3
2022 Quantum indistinguishability through exchangeability
abstract
Two particles are identical if all their intrinsic properties, such as spin and charge, are the same, meaning that no quantum experiment can distinguish them. In addition to the well known principles of quantum mechanics, understanding systems of identical particles requires a new postulate, the so called symmetrization postulate . In this work, we show that the postulate corresponds to exchangeability assessments for sets of observables (gambles) in a quantum experiment, when quantum mechanics is seen as a normative and algorithmic theory guiding an agent to assess her subjective beliefs represented as (coherent) sets of gambles. Finally, we show how sets of exchangeable observables (gambles) may be updated after a measurement and discuss the issue of defining entanglement for indistinguishable particle systems.
Alessio Benavoli, Alessandro Facchini, Marco Zaffalon
Int. J. Approx. Reason.1
2021 A Unified Framework for Closed-Form Nonparametric Regression, Classification, Preference and Mixed Problems with Skew Gaussian Processes
Alessio Benavoli, Dario Azzimonti, Dario Piga
DSAA1
2021 Bayesian Independence Test with Mixed-type Variables
abstract
A fundamental task in AI is to assess (in)dependence between mixed-type variables (text, image, sound). We propose a Bayesian kernelised correlation test of (in)dependence using a Dirichlet process model. The new measure of (in)dependence allows us to answer some fundamental questions: Based on data, are (mixed-type) variables independent? How likely is dependence/independence to hold? How high is the probability that two mixed-type variables are more than just weakly dependent? We theoretically show the properties of the approach, as well as algorithms for fast computation with it. We empirically demonstrate the effectiveness of the proposed method by analysing its performance and by comparing it with other frequentist and Bayesian approaches on a range of datasets and tasks with mixed-type variables.
Alessio Benavoli, Cassio P. de Campos
DSAA1
2021 Time Series Forecasting with Gaussian Processes Needs Priors
Giorgio Corani, Alessio Benavoli, Marco Zaffalon
ECML/PKDD (4)2
2021 Sparse Information Filter for Fast Gaussian Process Regression
Lucas Kania, Manuel Schürch, Dario Azzimonti, Alessio Benavoli
ECML/PKDD (3)4
2021 A unified framework for closed-form nonparametric regression, classification, preference and mixed problems with Skew Gaussian Processes
abstract
Abstract Skew-Gaussian Processes (SkewGPs) extend the multivariate Unified Skew-Normal distributions over finite dimensional vectors to distribution over functions. SkewGPs are more general and flexible than Gaussian processes, as SkewGPs may also represent asymmetric distributions. In a recent contribution, we showed that SkewGP and probit likelihood are conjugate, which allows us to compute the exact posterior for non-parametric binary classification and preference learning. In this paper, we generalize previous results and we prove that SkewGP is conjugate with both the normal and affine probit likelihood, and more in general, with their product. This allows us to (i) handle classification, preference, numeric and ordinal regression, and mixed problems in a unified framework; (ii) derive closed-form expression for the corresponding posterior distributions. We show empirically that the proposed framework based on SkewGP provides better performance than Gaussian processes in active learning and Bayesian (constrained) optimization. These two tasks are fundamental for design of experiments and in Data Science.
Alessio Benavoli, Dario Azzimonti, Dario Piga
Mach. Learn.1
2020 Skew Gaussian processes for classification
abstract
Abstract Gaussian processes (GPs) are distributions over functions, which provide a Bayesian nonparametric approach to regression and classification. In spite of their success, GPs have limited use in some applications, for example, in some cases a symmetric distribution with respect to its mean is an unreasonable model. This implies, for instance, that the mean and the median coincide, while the mean and median in an asymmetric (skewed) distribution can be different numbers. In this paper, we propose skew-Gaussian processes (SkewGPs) as a non-parametric prior over functions. A SkewGP extends the multivariateunified skew-normaldistribution over finite dimensional vectors to a stochastic processes. The SkewGP class of distributions includes GPs and, therefore, SkewGPs inherit all good properties of GPs and increase their flexibility by allowing asymmetry in the probabilistic model. By exploiting the fact that SkewGP and probit likelihood are conjugate model, we derive closed form expressions for the marginal likelihood and predictive distribution of this new nonparametric classifier. We verify empirically that the proposed SkewGP classifier provides a better performance than a GP classifier based on either Laplace’s method or expectation propagation.
Alessio Benavoli, Dario Azzimonti, Dario Piga
Mach. Learn.1
2019 Sum-of-squares for bounded rationality
Alessio Benavoli, Alessandro Facchini, Dario Piga, Marco Zaffalon
Int. J. Approx. Reason.1
2017 Introduction to the special issue on Bayesian Nonparametrics
Alessio Benavoli, Antonio Lijoi, Antonietta Mira
Int. J. Approx. Reason.1
2017 Time for a Change: a Tutorial for Comparing Multiple Classifiers Through Bayesian Analysis
abstract
The machine learning community adopted the use of null hypothesis significance testing (NHST) in order to ensure the statistical validity of results. Many scientific fields however realized the shortcomings of frequentist reasoning and in the most radical cases even banned its use in publications. We should do the same: just as we have embraced the Bayesian paradigm in the development of new machine learning methods, so we should also use it in the analysis of our own results. We argue for abandonment of NHST by exposing its fallacies and, more importantly, offer better---more sound and useful--- alternatives for it.
Alessio Benavoli, Giorgio Corani, Janez Demsar, Marco Zaffalon
J. Mach. Learn. Res.1
2017 Statistical comparison of classifiers through Bayesian hierarchical modelling
Giorgio Corani, Alessio Benavoli, Janez Demsar, Francesca Mangili, Marco Zaffalon
Mach. Learn.2
2016 Should We Really Use Post-Hoc Tests Based on Mean-Ranks?
abstract
The statistical comparison of multiple algorithms over multiple data sets is fundamental in machine learning. This is typically carried out by the Friedman test. When the Friedman test rejects the null hypothesis, multiple comparisons are carried out to establish which are the significant differences among algorithms. The multiple comparisons are usually performed using the mean-ranks test. The aim of this technical note is to discuss the inconsistencies of the mean-ranks post-hoc test with the goal of discouraging its use in machine learning as well as in medicine, psychology, etc.. We show that the outcome of the mean-ranks test depends on the pool of algorithms originally included in the experiment. In other words, the outcome of the comparison between algorithms $A$ and $B$ depends also on the performance of the other algorithms included in the original experiment. This can lead to paradoxical situations. For instance the difference between $A$ and $B$ could be declared significant if the pool comprises algorithms $C,D,E$ and not significant if the pool comprises algorithms $F,G,H$. To overcome these issues, we suggest instead to perform the multiple comparison using a test whose outcome only depends on the two algorithms being compared, such as the sign-test or the Wilcoxon signed-rank test.
Alessio Benavoli, Giorgio Corani, Francesca Mangili
J. Mach. Learn. Res.1
2015 Gaussian Processes for Bayesian hypothesis tests on regression functions
abstract
Gaussian processes have been used in different application domains such as classification, regression etc. In this paper we show that they can also be employed as a universal tool for developing a large variety of Bayesian statistical hypothesis tests for regression functions. In particular, we will use GPs for testing whether (i) two functions are equal; (ii) a function is monotone (even accounting for seasonality effects); (iii) a function is periodic; (iv) two functions are proportional. By simulation studies, we will show that, beside being more flexible, GP tests are also competitive in terms of performance with state-of-art algorithms.
Alessio Benavoli, Francesca Mangili
AISTATS1
2015 A Bayesian nonparametric procedure for comparing algorithms
abstract
A fundamental task in machine learning is to compare the performance of multiple algorithms. This is typically performed by frequentist tests (usually the Friedman test followed by a series of multiple pairwise comparisons). This implies dealing with null hypothesis significance tests and p-values, although the shortcomings of such methods are well known. First, we propose a nonparametric Bayesian version of the Friedman test using a Dirichlet process (DP) based prior. Our derivations show that, from a Bayesian perspective, the Friedman test is an inference for a multivariate mean based on an ellipsoid inclusion test. Second, we derive a joint procedure for the analysis of the multiple comparisons which accounts for their dependencies and which is based on the posterior probability computed through the DP. The proposed approach allows verifying the null hypothesis, not only rejecting it. Third, we apply our test to perform algorithms racing, i.e., the problem of identifying the best algorithm among a large set of candidates. We show by simulation that our approach is competitive both in terms of accuracy and speed in identifying the best algorithm.
Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon
ICML1
2015 Bayesian Hypothesis Testing in Machine Learning
Giorgio Corani, Alessio Benavoli, Francesca Mangili, Marco Zaffalon
ECML/PKDD (3)2
2015 New prior near-ignorance models on the simplex
Francesca Mangili, Alessio Benavoli
Int. J. Approx. Reason.2
2015 A Bayesian approach for comparing cross-validated algorithms on multiple data sets
Giorgio Corani, Alessio Benavoli
Mach. Learn.2
2014 A Bayesian Wilcoxon signed-rank test based on the Dirichlet process
abstract
Bayesian methods are ubiquitous in machine learning. Nevertheless, the analysis of empirical results is typically performed by frequentist tests. This implies dealing with null hypothesis significance tests and p-values, even though the shortcomings of such methods are well known. We propose a nonparametric Bayesian version of the Wilcoxon signed-rank test using a Dirichlet process (DP) based prior. We address in two different ways the problem of how to choose the infinite dimensional parameter that characterizes the DP. The proposed test has all the traditional strengths of the Bayesian approach; for instance, unlike the frequentist tests, it allows verifying the null hypothesis, not only rejecting it, and taking decision which minimize the expected loss. Moreover, one of the solutions proposed to model the infinitedimensional parameter of the DP, allows isolating instances in which the traditional frequentist test is guessing at random. We show results dealing with the comparison of two classifiers using real and simulated data.
Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, Fabrizio Ruggeri 0001
ICML1
2014 Belief function and multivalued mapping robustness in statistical estimation
Alessio Benavoli
Int. J. Approx. Reason.1
2014 Probabilistic Inference in Credal Networks: New Complexity Results
abstract
Credal networks are graph-based statistical models whose parameters take values in a set, instead of being sharply specified as in traditional statistical models (e.g., Bayesian networks). The computational complexity of inferences on such models depends on the irrelevance/independence concept adopted. In this paper, we study inferential complexity under the concepts of epistemic irrelevance and strong independence. We show that inferences under strong independence are NP-hard even in trees with binary variables except for a single ternary one. We prove that under epistemic irrelevance the polynomial-time complexity of inferences in credal trees is not likely to extend to more general models (e.g., singly connected topologies). These results clearly distinguish networks that admit efficient inferences and those where inferences are most likely hard, and settle several open questions regarding their computational complexity. We show that these results remain valid even if we disallow the use of zero probabilities. We also show that the computation of bounds on the probability of the future state in a hidden Markov model is the same whether we assume epistemic irrelevance or strong independence, and we prove a similar result for inference in naive Bayes structures. These inferential equivalences are important for practitioners, as hidden Markov models and naive Bayes structures are used in real applications of imprecise probability.
Denis Deratani Mauá, Cassio P. de Campos, Alessio Benavoli, Alessandro Antonucci 0001
J. Artif. Intell. Res.3
2013 Imprecise Hierarchical Dirichlet model with applications
Alessio Benavoli
FUSION1
2013 Set-membership PHD filter
Alessio Benavoli, Francesco Papi
FUSION1
2013 On the Complexity of Strong and Epistemic Credal Networks
Denis Deratani Mauá, Cassio P. de Campos, Alessio Benavoli, Alessandro Antonucci 0001
UAI3
2012 Pushing Kalman's idea to the extremes
Alessio Benavoli, Benjamin Noack
FUSION1
2011 Classification with imprecise likelihoods: A comparison of TBM, random set and imprecise probability approach
Alessio Benavoli, Branko Ristic 0001
FUSION1
2011 Inference with Multinomial Data: Why to Weaken the Prior Strength
abstract
This paper considers inference from multinomial data and addresses the problem of choosing the strength of the Dirichlet prior under a mean-squared error criterion. We compare the Maxi-mum Likelihood Estimator (MLE) and the most commonly used Bayesian estimators obtained by assuming a prior Dirichlet distribution with non-informative prior parameters, that is, the parameters of the Dirichlet are equal and altogether sum up to the so called strength of the prior. Under this criterion, MLE becomes more preferable than the Bayesian estimators at the increase of the number of categories k of the multinomial, because non-informative Bayesian estimators induce a region where they are dominant that quickly shrinks with the increase of k. This can be avoided if the strength of the prior is not kept constant but decreased with the number of categories. We argue that the strength should decrease at least k times faster than usual estimators do.
Cassio P. de Campos, Alessio Benavoli
IJCAI2
2010 Interval dominance based data association
Alessio Benavoli
FUSION1
2010 Restricting the IDM for Classification
Giorgio Corani, Alessio Benavoli
IPMU (1)2
2010 An aggregation framework based on coherent lower previsions: Application to Zadeh's paradox and sensor networks
Alessio Benavoli, Alessandro Antonucci 0001
Int. J. Approx. Reason.1
2009 Inference from Multinomial Data Based on a MLE-Dominance Criterion
Alessio Benavoli, Cassio P. de Campos
ECSQARU1
2009 Multiple model tracking by imprecise markov trees
Alessandro Antonucci 0001, Alessio Benavoli, Marco Zaffalon, Gert de Cooman, Filip Hermans
FUSION2
2009 Reliable hidden Markov model filtering through coherent lower previsions
Alessio Benavoli, Marco Zaffalon, Enrique Miranda 0001
FUSION1
2009 Fibonacci sequence, golden section, Kalman filter and optimal control
Alessio Benavoli, Luigi Chisci, Alfonso Farina
Signal Process.1
2008 Modelling uncertain implication rules in evidence theory
Alessio Benavoli, Luigi Chisci, Alfonso Farina, Branko Ristic 0001
FUSION1
2007 An approach to threat assessment based on evidential networks
abstract
The paper develops an information fusion system that aims at supporting a commander's decision making by providing an assessment of threat, that is an estimate of the extent to which an enemy platform poses a threat based on evidence about its intent and capability. Threat is modelled in the framework of the valuation-based system (VBS), by a network of entities and relationships between them. The uncertainties in the relationships are represented by belief functions as defined in the theory of evidence. Hence the resulting network for reasoning is referred to as an evidential network. Local computations in the evidential network are carried out by inward propagation on the underlying joint binary tree. This allows the dynamic nature of the external evidence, which drives the evidential network, to be taken into account by recomputing only the affected paths in the joint binary tree.
Alessio Benavoli, Branko Ristic 0001, Alfonso Farina, Martin Oxenham, Luigi Chisci
FUSION1