VLDB 2026 Research / reviewers in the wild / expert
Michele Donini
dblp:149/0239
· DBLP profile ↗
37ranked-venue papers
5as first author
10since 2021 · last 2024
0000-0002-9769-3899ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Explaining Probabilistic Models with Distributional ValuesabstractA large branch of explainable machine learning is grounded in cooperative game theory. However, research indicates that game-theoretic explanations may mislead or be hard to interpret. We argue that often there is a critical mismatch between what one wishes to explain (e.g. the output of a classifier) and what current methods such as SHAP explain (e.g. the scalar probability of a class). This paper addresses such gap for probabilistic models by generalising cooperative games and value operators. We introduce the *distributional values*, random variables that track changes in the model output (e.g. flipping of the predicted class) and derive their analytic expressions for games with Gaussian, Bernoulli and Categorical payoffs. We further establish several characterising properties, and show that our framework provides fine-grained and insightful explanations with case studies on vision and language models. Luca Franceschi 0001, Michele Donini, Cédric Archambeau, Matthias W. Seeger |
ICML | 2 |
| 2024 | Fortuna: A Library for Uncertainty Quantification in Deep LearningabstractWe present Fortuna, an open-source library for uncertainty quantification in deep learning. Fortuna supports a range of calibration techniques, such as conformal prediction that can be applied to any trained neural network to generate reliable uncertainty estimates, and scalable Bayesian inference methods that can be applied to deep neural networks trained from scratch for improved uncertainty quantification and accuracy. By providing a coherent framework for advanced uncertainty quantification methods, Fortuna simplifies the process of benchmarking and helps practitioners build robust AI systems. Gianluca Detommaso, Alberto Gasparin, Michele Donini, Matthias W. Seeger, Andrew Gordon Wilson, Cédric Archambeau |
J. Mach. Learn. Res. | 3 |
| 2023 | Model AI Assignments 2023abstractThe Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of six AI assignments from the 2023 session that are easily adoptable, playfully engaging, and flexible for a variety of instructor needs. Assignment specifications and supporting resources may be found at http://modelai.gettysburg.edu . Todd W. Neller, Raechel Walker, Olivia Dias, Zeynep Yalcin, Cynthia Breazeal, Matthew E. Taylor, Michele Donini, Erin Talvitie, Charlie Pilgrim, Paolo Turrini, James Maher, Matthew Boutell, Justin Wilson, Narges Norouzi, Jonathan Scott |
AAAI | 7 |
| 2023 | Efficient fair PCA for fair representation learningabstractWe revisit the problem of fair principal component analysis (PCA), where the goal is to learn the best low-rank linear approximation of the data that obfuscates demographic information. We propose a conceptually simple approach that allows for an analytic solution similar to standard PCA and can be kernelized. Our methods have the same complexity as standard PCA, or kernel PCA, and run much faster than existing methods for fair PCA based on semidefinite programming or manifold optimization, while achieving similar results. Matthäus Kleindessner, Michele Donini, Chris Russell 0001, Muhammad Bilal Zafar |
AISTATS | 2 |
| 2022 | Amazon SageMaker Model Monitor: A System for Real-Time Insights into Deployed Machine Learning ModelsabstractWith the increasing adoption of machine learning (ML) models and systems in high-stakes settings across different industries, guaranteeing a model's performance after deployment has become crucial. Monitoring models in production is a critical aspect of ensuring their continued performance and reliability. We present Amazon SageMaker Model Monitor, a fully managed service that continuously monitors the quality of machine learning models hosted on Amazon SageMaker. Our system automatically detects data, concept, bias, and feature attribution drift in models in real-time and provides alerts so that model owners can take corrective actions and thereby maintain high quality models. We describe the key requirements obtained from customers, system design and architecture, and methodology for detecting different types of drift. Further, we provide quantitative evaluations followed by use cases, insights, and lessons learned from more than two years of production deployment. David Nigenda, Zohar S. Karnin, Muhammad Bilal Zafar, Raghu Ramesha, Alan Tan, Michele Donini, Krishnaram Kenthapadi |
KDD | 6 |
| 2022 | Deep fair models for complex data: Graphs labeling and explainable face recognition
Danilo Franco, Nicolò Navarin, Michele Donini, Davide Anguita, Luca Oneto |
Neurocomputing | 3 |
| 2021 | Fair Bayesian OptimizationabstractGiven the increasing importance of machine learning (ML) in our lives, several algorithmic fairness techniques have been proposed to mitigate biases in the outcomes of the ML models. However, most of these techniques are specialized to cater to a single family of ML models and a specific definition of fairness, limiting their adaptibility in practice. We introduce a general constrained Bayesian optimization (BO) framework to optimize the performance of any ML model while enforcing one or multiple fairness constraints. BO is a model-agnostic optimization method that has been successfully applied to automatically tune the hyperparameters of ML models. We apply BO with fairness constraints to a range of popular models, including random forests, gradient boosting, and neural networks, showing that we can obtain accurate and fair solutions by acting solely on the hyperparameters. We also show empirically that our approach is competitive with specialized techniques that enforce model-specific fairness constraints, and outperforms preprocessing methods that learn fair representations of the input data. Moreover, our method can be used in synergy with such specialized fairness techniques to tune their hyperparameters. Finally, we study the relationship between fairness and the hyperparameters selected by BO. We observe a correlation between regularization and unbiased models, explaining why acting on the hyperparameters leads to ML models that generalize well and are fair. Valerio Perrone, Michele Donini, Muhammad Bilal Zafar, Robin Schmucker, Krishnaram Kenthapadi, Cédric Archambeau |
AIES | 2 |
| 2021 | Amazon SageMaker Clarify: Machine Learning Bias Detection and Explainability in the CloudabstractUnderstanding the predictions made by machine learning (ML) models and their potential biases remains a challenging and labor-intensive task that depends on the application, the dataset, and the specific model. We present Amazon SageMaker Clarify, an explainability feature for Amazon SageMaker that launched in December 2020, providing insights into data and ML models by identifying biases and explaining predictions. It is deeply integrated into Amazon SageMaker, a fully managed service that enables data scientists and developers to build, train, and deploy ML models at any scale. Clarify supports bias detection and feature importance computation across the ML lifecycle, during data preparation, model evaluation, and post-deployment monitoring. We outline the desiderata derived from customer input, the modular architecture, and the methodology for bias and explanation computations. Further, we describe the technical challenges encountered and the tradeoffs we had to make. For illustration, we discuss two customer use cases. We present our deployment results including qualitative customer feedback and a quantitative evaluation. Finally, we summarize lessons learned, and discuss best practices for the successful adoption of fairness and explanation tools in practice. Michaela Hardt, Xiaoyi Cheng, Michele Donini, Jason Gelman, Satish Gollaprolu, John He, Pedro Larroy, Nick McCarthy, Ashish Rathi, Scott Rees, Amaresh Ankit Siva, ErhYuan Tsai, Keerthan Vasist, Pinar Yilmaz, Muhammad Bilal Zafar, Sanjiv Das, Kevin Haas, Tyler Hill, Krishnaram Kenthapadi |
KDD | 4 |
| 2021 | Amazon SageMaker Automatic Model Tuning: Scalable Gradient-Free OptimizationabstractTuning complex machine learning systems is challenging. Machine learning typically requires to set hyperparameters, be it regularization, architecture, or optimization parameters, whose tuning is critical to achieve good predictive performance. To democratize access to machine learning systems, it is essential to automate the tuning. This paper presents Amazon SageMaker Automatic Model Tuning (AMT), a fully managed system for gradient-free optimization at scale. AMT finds the best version of a trained machine learning model by repeatedly evaluating it with different hyperparameter configurations. It leverages either random search or Bayesian optimization to choose the hyperparameter values resulting in the best model, as measured by the metric chosen by the user. AMT can be used with built-in algorithms, custom algorithms, and Amazon SageMaker pre-built containers for machine learning frameworks. We discuss the core functionality, system architecture, our design principles, and lessons learned. We also describe more advanced features of AMT, such as automated early stopping and warm-starting, showing in experiments their benefits to users. Valerio Perrone, Huibin Shen, Aida Zolic, Iaroslav Shcherbatyi, Amr Ahmed 0004, Tanya Bansal, Michele Donini, Fela Winkelmolen, Rodolphe Jenatton, Jean Baptiste Faddoul, Barbara Pogorzelska, Miroslav Miladinovic, Krishnaram Kenthapadi, Matthias W. Seeger, Cédric Archambeau |
KDD | 7 |
| 2021 | Voting with random classifiers (VORACE): theoretical and experimental analysisabstractAbstract In many machine learning scenarios, looking for the best classifier that fits a particular dataset can be very costly in terms of time and resources. Moreover, it can require deep knowledge of the specific domain. We propose a new technique which does not require profound expertise in the domain and avoids the commonly used strategy of hyper-parameter tuning and model selection. Our method is an innovative ensemble technique that uses voting rules over a set of randomly-generated classifiers. Given a new input sample, we interpret the output of each classifier as a ranking over the set of possible classes. We then aggregate these output rankings using a voting rule, which treats them as preferences over the classes. We show that our approach obtains good results compared to the state-of-the-art, both providing a theoretical analysis and an empirical evaluation of the approach on several datasets. Cristina Cornelio, Michele Donini, Andrea Loreggia, Maria Silvia Pini, Francesca Rossi 0001 |
Auton. Agents Multi Agent Syst. | 2 |
| 2020 | Learning Fair and Transferable Representations with Theoretical GuaranteesabstractDeveloping learning methods which do not discriminate subgroups in the population is the central goal of algorithmic fairness. One way to reach this goal is by modifying the data representation in order to satisfy prescribed fairness constraints. This allows to reuse the same representation in other context (tasks) without discriminate subgroups. In this work we measure fairness according to demographic parity, requiring the probability of the possible model decisions to be independent of the sensitive information. We argue that the goal of imposing demographic parity can be substantially facilitated within a multi-task learning setting. We leverage task similarities by encouraging a shared fair representation across the tasks via low rank matrix factorization. We derive learning bounds establishing that the learned representation transfers well to novel tasks both in terms of prediction performance and fairness metrics. We present experiments on three real world datasets, showing that the proposed method outperforms state-of-the-art approaches by a significant margin. Luca Oneto, Michele Donini, Massimiliano Pontil, Andreas Maurer |
DSAA | 2 |
| 2020 | Learning Deep Fair Graph Neural Networks
Luca Oneto, Nicolò Navarin, Michele Donini |
ESANN | 3 |
| 2020 | Marthe: Scheduling the Learning Rate Via Online HypergradientsabstractWe study the problem of fitting task-specific learning rate schedules from the perspective of hyperparameter optimization, aiming at good generalization. We describe the structure of the gradient of a validation error w.r.t. the learning rate schedule -- the hypergradient. Based on this, we introduce MARTHE, a novel online algorithm guided by cheap approximations of the hypergradient that uses past information from the optimization trajectory to simulate future behaviour. It interpolates between two recent techniques, RTHO (Franceschi et al., 2017) and HD (Baydin et al. 2018), and is able to produce learning rate schedules that are more stable leading to models that generalize better. Michele Donini, Luca Franceschi 0001, Orchid Majumder, Massimiliano Pontil, Paolo Frasconi |
IJCAI | 1 |
| 2020 | General Fair Empirical Risk MinimizationabstractWe tackle the problem of algorithmic fairness, where the goal is to avoid the unfairly influence of sensitive information, in the general context of regression with possible continuous sensitive attributes. We extend the framework of fair empirical risk minimization of [1] to this general scenario, covering in this way the whole standard supervised learning setting. Our generalized fairness measure reduces to well known notions of fairness available in literature. We derive learning guarantees for our method, that imply in particular its statistical consistency, both in terms of the risk and the fairness measure. We then specialize our approach to kernel methods and propose a convex fair estimator in that setting. We test the estimator on a commonly used benchmark dataset (Communities and Crime) and on a new dataset collected at the University of Genoa1, containing the information of the academic career of five thousand students. The latter dataset provides a challenging real case scenario of unfair behaviour of standard regression methods that benefits from our methodology. The experimental results show that our estimator is effective at mitigating the trade-off between accuracy and fairness requirements. Luca Oneto, Michele Donini, Massimiliano Pontil |
IJCNN | 2 |
| 2020 | Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningabstractDeveloping learning methods which do not discriminate subgroups in the population is a central goal of algorithmic fairness. One way to reach this goal is by modifying the data representation in order to meet certain fairness constraints. In this work we measure fairness according to demographic parity. This requires the probability of the possible model decisions to be independent of the sensitive information. We argue that the goal of imposing demographic parity can be substantially facilitated within a multitask learning setting. We present a method for learning a shared fair representation across multiple tasks, by means of different new constraints based on MMD and Sinkhorn Divergences. We derive learning bounds establishing that the learned representation transfers well to novel tasks. We present experiments on three real world datasets, showing that the proposed method outperforms state-of-the-art approaches by a significant margin. Luca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto, Andreas Maurer, Massimiliano Pontil |
NeurIPS | 2 |
| 2020 | Randomized learning and generalization of fair and private classifiers: From PAC-Bayes to stability and differential privacy
Luca Oneto, Michele Donini, Massimiliano Pontil, John Shawe-Taylor |
Neurocomputing | 2 |
| 2019 | Taking Advantage of Multitask Learning for Fair ClassificationabstractA central goal of algorithmic fairness is to reduce bias in automated decision making. An unavoidable tension exists between accuracy gains obtained by using sensitive information as part of a statistical model, and any commitment to protect these characteristics. Often, due to biases present in the data, using the sensitive information in the functional form of a classifier improves classification accuracy. In this paper we show how it is possible to get the best of both worlds: optimize model accuracy and fairness without explicitly using the sensitive feature in the functional form of the model, thereby treating different individuals equally. Our method is based on two key ideas. On the one hand, we propose to use Multitask Learning (MTL), enhanced with fairness constraints, to jointly learn group specific classifiers that leverage information between sensitive groups. On the other hand, since learning group specific models might not be permitted, we propose to first predict the sensitive features by any learning method and then to use the predicted sensitive feature to train MTL with fairness constraints. This enables us to tackle fairness with a three-pronged approach, that is, by increasing accuracy on each group, enforcing measures of fairness during training, and protecting sensitive information during testing. Experimental results on two real datasets support our proposal, showing substantial improvements in both accuracy and fairness. Luca Oneto, Michele Donini, Amon Elders, Massimiliano Pontil |
AIES | 2 |
| 2019 | PAC-Bayes and Fairness: Risk and Fairness Bounds on Distribution Dependent Fair Priors
Luca Oneto, Michele Donini, Massimiliano Pontil |
ESANN | 2 |
| 2018 | Emerging trends in machine learning: beyond conventional methods and data
Luca Oneto, Nicolò Navarin, Michele Donini, Davide Anguita |
ESANN | 3 |
| 2018 | Empirical Risk Minimization Under Fairness ConstraintsabstractWe address the problem of algorithmic fairness: ensuring that sensitive information does not unfairly influence the outcome of a classifier. We present an approach based on empirical risk minimization, which incorporates a fairness constraint into the learning problem. It encourages the conditional risk of the learned classifier to be approximately constant with respect to the sensitive variable. We derive both risk and fairness bounds that support the statistical consistency of our methodology. We specify our approach to kernel methods and observe that the fairness requirement implies an orthogonality constraint which can be easily added to these methods. We further observe that for linear models the constraint translates into a simple data preprocessing step. Experiments indicate that the method is empirically effective and performs favorably against state-of-the-art approaches. Michele Donini, Luca Oneto, Shai Ben-David, John Shawe-Taylor, Massimiliano Pontil |
NeurIPS | 1 |
| 2018 | Scuba: scalable kernel-based gene prioritizationabstractBACKGROUND: The uncovering of genes linked to human diseases is a pressing challenge in molecular biology and precision medicine. This task is often hindered by the large number of candidate genes and by the heterogeneity of the available information. Computational methods for the prioritization of candidate genes can help to cope with these problems. In particular, kernel-based methods are a powerful resource for the integration of heterogeneous biological knowledge, however, their practical implementation is often precluded by their limited scalability. RESULTS: We propose Scuba, a scalable kernel-based method for gene prioritization. It implements a novel multiple kernel learning approach, based on a semi-supervised perspective and on the optimization of the margin distribution. Scuba is optimized to cope with strongly unbalanced settings where known disease genes are few and large scale predictions are required. Importantly, it is able to efficiently deal both with a large amount of candidate genes and with an arbitrary number of data sources. As a direct consequence of scalability, Scuba integrates also a new efficient strategy to select optimal kernel parameters for each data source. We performed cross-validation experiments and simulated a realistic usage setting, showing that Scuba outperforms a wide range of state-of-the-art methods. CONCLUSIONS: Scuba achieves state-of-the-art performance and has enhanced scalability compared to existing kernel-based approaches for genomic data. This method can be useful to prioritize candidate genes, particularly when their number is large or when input data is highly heterogeneous. The code is freely available at https://github.com/gzampieri/Scuba . Guido Zampieri, Michele Donini, Nicolò Navarin, Fabio Aiolli, Alessandro Sperduti, Giorgio Valle |
BMC Bioinform. | 3 |
| 2018 | Learning With Kernels: A Local Rademacher Complexity-Based Analysis With Application to Graph KernelsabstractWhen dealing with kernel methods, one has to decide which kernel and which values for the hyperparameters to use. Resampling techniques can address this issue but these procedures are time-consuming. This problem is particularly challenging when dealing with structured data, in particular with graphs, since several kernels for graph data have been proposed in literature, but no clear relationship among them in terms of learning properties is defined. In these cases, exhaustive search seems to be the only reasonable approach. Recently, the global Rademacher complexity (RC) and local Rademacher complexity (LRC), two powerful measures of the complexity of a hypothesis space, have shown to be suited for studying kernels properties. In particular, the LRC is able to bound the generalization error of an hypothesis chosen in a space by disregarding those ones which will not be taken into account by any learning procedure because of their high error. In this paper, we show a new approach to efficiently bound the RC of the space induced by a kernel, since its exact computation is an NP-Hard problem. Then we show for the first time that RC can be used to estimate the accuracy and expressivity of different graph kernels under different parameter configurations. The authors' claims are supported by experimental results on several real-world graph data sets. Luca Oneto, Nicolò Navarin, Michele Donini, Sandro Ridella, Alessandro Sperduti, Fabio Aiolli, Davide Anguita |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Fast hyperparameter selection for graph kernels via subsampling and multiple kernel learning
Michele Donini, Nicolò Navarin, Ivano Lauriola, Fabio Aiolli, Fabrizio Costa |
ESANN | 1 |
| 2017 | Learning dot-product polynomials for multiclass problems
Ivano Lauriola, Michele Donini, Fabio Aiolli |
ESANN | 2 |
| 2017 | Forward and Reverse Gradient-Based Hyperparameter OptimizationabstractWe study two procedures (reverse-mode and forward-mode) for computing the gradient of the validation error with respect to the hyperparameters of any iterative learning algorithm such as stochastic gradient descent. These procedures mirror two ways of computing gradients for recurrent neural networks and have different trade-offs in terms of running time and space requirements. Our formulation of the reverse-mode procedure is linked to previous work by Maclaurin et al (2015) but does not require reversible dynamics. Additionally, we explore the use of constraints on the hyperparameters. The forward-mode procedure is suitable for real-time hyperparameter updates, which may significantly speedup hyperparameter optimization on large datasets. We present a series of experiments on image and phone classification tasks. In the second task, previous gradient-based approaches are prohibitive. We show that our real-time algorithm yields state-of-the-art results in affordable time. Luca Franceschi 0001, Michele Donini, Paolo Frasconi, Massimiliano Pontil |
ICML | 2 |
| 2017 | A Speaker Adaptive DNN Training Approach for Speaker-Independent Acoustic InversionabstractInternational audience Leonardo Badino, Luca Franceschi 0001, Raman Arora, Michele Donini, Massimiliano Pontil |
INTERSPEECH | 4 |
| 2017 | Measuring the expressivity of graph kernels through Statistical Learning Theory
Luca Oneto, Nicolò Navarin, Michele Donini, Alessandro Sperduti, Fabio Aiolli, Davide Anguita |
Neurocomputing | 3 |
| 2017 | Learning deep kernels in the space of dot product polynomials
Michele Donini, Fabio Aiolli |
Mach. Learn. | 1 |
| 2016 | Advances in Learning with Kernels: Theory and Practice in a World of growing Constraints
Luca Oneto, Nicolò Navarin, Michele Donini, Fabio Aiolli, Davide Anguita |
ESANN | 3 |
| 2016 | Measuring the Expressivity of Graph Kernels through the Rademacher Complexity
Luca Oneto, Nicolò Navarin, Michele Donini, Alessandro Sperduti, Fabio Aiolli, Davide Anguita |
ESANN | 3 |
| 2016 | Distributed variance regularized Multitask LearningabstractPast research on Multitask Learning (MTL) has focused mainly on devising adequate regularizers and less on their scalability. In this paper, we present a method to scale up MTL methods which penalize the variance of the task weight vectors. The method builds upon the alternating direction method of multipliers to decouple the variance regularizer. It can be efficiently implemented by a distributed algorithm, in which the tasks are first independently solved and subsequently corrected to pool information from other tasks. We show that the method works well in practice and convergences in few distributed iterations. Furthermore, we empirically observe that the number of iterations is nearly independent of the number of tasks, yielding a computational gain of O(T) over standard solvers. We also present experiments on a large URL classification dataset, which is challenging both in terms of volume of data points and dimensionality. Our results confirm that MTL can obtain superior performance over either learning a common model or independent task learning. Michele Donini, David Martínez-Rego, Martin Goodson, John Shawe-Taylor, Massimiliano Pontil |
IJCNN | 1 |
| 2016 | Stairstep recognition and counting in a serious Game for increasing users' physical activity
Matteo Ciman, Michele Donini, Ombretta Gaggi, Fabio Aiolli |
Pers. Ubiquitous Comput. | 2 |
| 2015 | Feature and kernel learning
Verónica Bolón-Canedo, Michele Donini, Fabio Aiolli |
ESANN | 2 |
| 2015 | EasyMKL: a scalable multiple kernel learning algorithm
Fabio Aiolli, Michele Donini |
Neurocomputing | 2 |
| 2014 | Easy multiple kernel learning
Fabio Aiolli, Michele Donini |
ESANN | 2 |
| 2014 | Learning Anisotropic RBF Kernels
Fabio Aiolli, Michele Donini |
ICANN | 2 |
| 2014 | ClimbTheWorld: real-time stairstep counting to increase physical activityabstractThe increasing number of people that are overweight due to a sedentary life requires persuasive strategies to convince people to change their behaviors. In this paper, we present a machine learning based technique to recognize and count stairsteps when a person climbs or descends stairs. This techn Fabio Aiolli, Matteo Ciman, Michele Donini, Ombretta Gaggi |
MobiQuitous | 3 |